iSCSI Hub
Flat isometric illustration of a black disk array with rows of pink-lit drive bays, standing on a dashed magenta grid against a deep indigo background.
troubleshooting

iSCSI Troubleshooting: Login, Timeout and Path Errors

A layered method for iSCSI faults: portal reachability, discovery, login, ACL and CHAP mismatches, MTU stalls, and the timers that decide when I/O fails.

By iSCSI Hub Editorial · ·Updated · 8 min read

iSCSI faults are diagnosed badly because the symptom always arrives at the wrong layer. An MTU mismatch on a switch presents as a database timing out. A CHAP typo presents as a disk that will not appear. A dead path presents, minutes later, as a filesystem going read-only. Working from the symptom backwards wastes time; working up the stack in a fixed order does not.

Start with the exact error, then check reachability, discovery and authentication before changing storage or recovery settings. The messages below come from the open-iscsi source; platform versions can add target names and further diagnostic text.

iscsiadm: initiator reported error (8 - connection timed out)

iscsiadm: initiator reported error (8 - connection timed out)

The upstream error table assigns code 8 to a connection timeout. This identifies a failed connection attempt, not a damaged filesystem. Check the portal address, storage-interface route, target listener and firewall policy before extending timeouts.

Before anything protocol-specific, confirm TCP reachability to the target portal on port 3260 from the host, on the interface that is supposed to carry storage traffic. Not from a management interface, and not by pinging the array’s management address. A surprising share of “iSCSI is broken” turns out to be a host that has one route to the storage subnet and it goes out the wrong NIC.

Microsoft’s iSCSI troubleshooting checklist opens in the same place, asking whether iSCSI, management and client networks are segregated and correctly routed, and whether MTU, VLAN, jumbo frame and flow control settings are consistent.

If the port does not answer, stop. Nothing further in this list is relevant.

no portals found

iscsiadm: No portals found

The open-iscsi discovery code emits this message when its result list contains no portal records. Read preceding messages: failed connection attempts and authentication diagnostics need different fixes from an empty discovery response.

Confirm that discovery uses the target’s storage address and the intended initiator interface. On the target, check that the portal is enabled, the intended target is advertised to this initiator, and any discovery CHAP settings match. Finding no portal does not prove that no LUN exists; it means this discovery operation supplied no usable portal records. Once discovery returns a target IQN and address, check login separately.

Login authentication failed

Login authentication failed

This is the diagnostic prefix used by open-iscsi’s login authentication checks; the full message can include the target IQN and a reason. Compare CHAP usernames, secrets and authentication direction on both ends without printing credentials into a diagnostic log.

These are two separate exchanges and separating them is the most valuable single diagnostic in iSCSI.

Discovery is a SendTargets request to the portal. If it returns a target list, the network path and the target service are both alive. Login is the session establishment against a specific target, and it is where identity and authentication are evaluated.

If the advertised login portal is reachable, check these three configuration differences:

  • ACL mismatch. The target does not have this initiator’s IQN mapped to any LUN. Read the initiator name from the client rather than from documentation: /etc/iscsi/initiatorname.iscsi on Linux, the Configuration tab of the iSCSI Initiator control panel on Windows. A single transposed character produces exactly this symptom.
  • CHAP mismatch. Check for a wrong secret, one-way CHAP on one side and mutual CHAP on the other, or a secret that violates the implementation’s length rules. RFC 7143 section 9.2.1 requires support for secrets up to 128 random bits and requires IPsec when a secret contains fewer than 96 random bits. These are randomness and protection requirements, not vendor character-length limits: character count alone does not establish randomness. Check the initiator and target’s documented character restrictions separately.
  • Discovery-time versus session-time authentication. open-iscsi treats these as separate settings, discovery.sendtargets.auth.* and node.session.auth.*, both defaulting to no authentication. Configuring one and not the other yields a target that lists fine and refuses to log in.

Where the target enforces CHAP on discovery too, the failure moves one step earlier and discovery itself returns nothing, which is worth knowing before concluding the portal is unreachable.

Login succeeds but the disk misbehaves

Once a session is up, the remaining faults are about data, not identity.

MTU mismatch is the signature failure here. Small packets get through, so discovery and login both succeed, and then anything that fills a frame stalls or retries. If jumbo frames are enabled anywhere in the path they must be enabled everywhere in it, including the switch, and Microsoft’s guidance repeats the point in both its checklist and its resolution steps: make sure MTUs and jumbo frames are consistent end to end. The cheapest test is to revert every hop to 1500 and see whether the problem disappears.

One disk per path means MPIO is not claiming the device. Multiple sessions to the same LUN without a multipath layer above them present as several independent disks with identical contents, and writing through two of them corrupts the volume. Aggregation is the whole job of that layer: Microsoft’s description of the initiator’s Devices tab is that with MPIO the iSCSI Initiator can log in on multiple sessions to the same target and aggregate the duplicate devices into a single device exposed to Windows, each session using different network adapters, network infrastructure and target ports. Without it, nothing merges them. On Windows the checks are Install-WindowsFeature Multipath-IO for the feature itself and mpclaim -s -d to see whether the Microsoft DSM has claimed the disks; on Linux the equivalent is that multipath -ll shows one map with several paths rather than nothing at all.

For the configuration sequence behind the Windows checks, see Windows Server target and MPIO setup.

Paths sharing one physical dependency need investigation. Microsoft’s Windows MPIO guidance calls for each redundant iSCSI connection to use a different network adapter, and a system that detects only one path needs its connections checked. Choose subnet layout according to the host platform and the storage vendor’s supported topology; a shared subnet alone does not diagnose failed multipathing. Verify the source adapter, network infrastructure and target port for each session, then test failover. The reasoning is set out in iSCSI fundamentals: targets, LUNs and multipathing.

The timers that decide when I/O fails

This is where “it recovered but the application died anyway” gets explained, and the defaults are worth memorising.

open-iscsi pings each session with NOP-Out requests. The shipped iscsid.conf sets node.conn[0].timeo.noop_out_interval = 5 and node.conn[0].timeo.noop_out_timeout = 5, both documented as five seconds. When a NOP-Out times out, the iSCSI layer fails the running commands and instructs the SCSI layer to requeue them.

The session then has a recovery wait before failure is passed upward. The upstream node.session.timeo.replacement_timeout = 120 sets the iSCSI replacement timeout, but this value alone does not determine the effective wait on a DM Multipath device.

For DM Multipath, inspect fast_io_fail_tmo in the active multipath configuration first. Red Hat’s link-loss guidance states that fast_io_fail_tmo takes precedence over replacement_timeout and advises against using replacement_timeout to override the session’s recovery_tmo on DM Multipath devices. Review the effective recovery timeout, multipath queueing policy and application timeouts together before changing them.

The same guide gives a 15 to 20 second replacement_timeout example with queue_if_no_path enabled. That example applies only where replacement_timeout governs recovery; it is not a universal DM Multipath recommendation and does not override fast_io_fail_tmo. A configured value of 120 alone does not establish that multipath I/O will stall for two minutes. Validate failover and all-paths-lost behavior for the actual configuration during a maintenance window.

On Windows the equivalent lever is the disk timeout. Microsoft’s guidance for surprise-removal and failover problems includes setting TimeOutValue under HKLM\SYSTEM\CurrentControlSet\Services\disk to a larger number, such as 179. Change it deliberately: a longer disk timeout hides transient path loss from applications, and also hides real path loss for the same interval.

The other open-iscsi default worth knowing is node.startup = manual. A session that was working perfectly and vanishes after a reboot is usually this, not a fault.

Reading the Windows event log

Microsoft’s troubleshooting article lists the events that matter, and they are far more specific than the generic disk errors that surround them:

  • Event ID 157, “Disk X has been surprise removed”, is the classic signature of a path or session dropping under an active volume.
  • Event IDs 9, 20, 27, 39 and 153 cover the iSCSI side, including “Target did not respond”, “Initiator failed to connect” and retried I/O at a logical block address.
  • The documented causes are network instability, MPIO configuration errors, adapters or NIC teams not being ready when the iSCSI service starts so ports cannot bind, mismatched VLAN or MTU or jumbo settings, outdated firmware and drivers, and resource exhaustion on the array.

For state rather than history, Microsoft points at Get-IscsiConnection, Get-IscsiSession and Get-MSDSMAutomaticClaimSettings to gather path and session status, and Get-Disk and Get-PhysicalDisk to review the resulting disk mappings.

Two items in that list are easy to miss and expensive to rediscover. First, applications and scripts must not rely on disk numbers, because path failovers can change them. Second, LBFO NIC teaming is deprecated for Hyper-V deployments as of Windows Server 2022, with switch embedded teaming as the replacement, so a teaming configuration inherited from an older build is a legitimate suspect rather than a stable baseline.

When the volume itself is damaged

If a volume has gone RAW or a filesystem is reporting checksum errors, the protocol layer is no longer the problem and the priority order changes: back up the affected volume first, repair second. Microsoft’s sequence is explicit about that ordering, with chkdsk /f and chkdsk /r for NTFS and refsutil salvage for ReFS, and with the log and recovery folders required to be on a different volume from the damaged one.

Two causes deserve checking before the repair, because both will recreate the damage: a LUN attached read-write by more than one host without a cluster-aware filesystem, and a thin-provisioned LUN whose backing store filled up. Neither is a filesystem bug. The first is an architectural mistake described in iSCSI vs NFS vs SMB, where the comparison of ownership models makes clear why block storage arbitrates nothing.

A short checklist

  1. TCP 3260 answers from the storage interface, not the management one.
  2. Discovery returns a target list.
  3. The target’s ACL contains the initiator’s actual IQN, copied from the client.
  4. CHAP is configured on the same side, in the same direction, at both discovery and session scope.
  5. MTU is identical on host, switch and target, or 1500 everywhere.
  6. Each path uses the intended adapter and target port within the platform’s supported network topology.
  7. The host shows one multipath device, not one disk per path.
  8. Effective recovery timers and queueing policy match the topology and application requirements; inspect fast_io_fail_tmo first for DM Multipath.
  9. Node startup is automatic if the LUN is expected to survive a reboot.

For setup decisions on NAS targets, see Unraid, TrueNAS and Synology iSCSI setup. To estimate whether the intended paths can carry the workload, use the iSCSI Throughput and MPIO Calculator.

Sources

  1. open-iscsi: error numbers and initiator error messages
  2. open-iscsi: login authentication diagnostics
  3. open-iscsi: SendTargets discovery and empty portal results
  4. open-iscsi: default iscsid.conf with documented timer values
  5. Red Hat Enterprise Linux: Modifying link loss behavior
  6. Microsoft Learn: iSCSI storage connectivity troubleshooting guidance
  7. Microsoft Learn: Multipath I/O (MPIO) troubleshooting guidance
  8. Microsoft iSCSI Initiator: Devices and MPIO (archived Windows Server documentation)
  9. RFC 7143: Internet Small Computer System Interface (iSCSI) Protocol (Consolidated)

Related