Skip to main content

Why I Chose iSCSI for the Back-End Storage Layer

Based on a LinkedIn post originally published on 4 February 2026

For the back-end storage layer of my home lab, I deliberately chose iSCSI instead of attaching every physical storage device directly to one host.

This was not primarily a convenience decision.

It was driven by architecture, scalability, resource isolation, maintenance and the need to keep failure behaviour understandable as the platform grows.

In this design, storage devices are distributed across several back-end nodes. Those nodes expose block devices over iSCSI to HG000009, the front-end data store.

HG000009 then aggregates and manages the storage through ZFS.

The architecture can be represented conceptually as:

Physical storage devices
        ↓
Back-end storage nodes
        ↓
iSCSI targets
        ↓
Dedicated storage network
        ↓
HG000009 iSCSI initiator
        ↓
ZFS storage control plane
        ↓
Client-facing storage services

iSCSI is therefore more than a transport mechanism in this design.

It creates a deliberate boundary between raw storage and the system responsible for organising, protecting and presenting the data.

Why not attach every device locally?

The simplest initial design would have been to attach every storage enclosure directly to HG000009.

For a small number of devices, that approach can work.

As the number of disks grows, several problems emerge.

USB topology and practical limits

Linux can manage many USB devices, but the USB architecture has finite addressing and topology limits.

The theoretical number is less important than the practical reality.

Real environments must account for:

  • controllers;

  • hubs;

  • hub depth;

  • enclosure bridges;

  • power delivery;

  • cabling;

  • device resets;

  • bandwidth sharing;

  • driver behaviour;

  • and the reliability of USB-to-storage adapters.

An enclosure may expose more than one logical USB device. Hubs also consume positions in the topology. A design that appears to have room for many additional disks on paper may become unreliable long before reaching a theoretical maximum.

Designing a storage platform around the maximum number of devices that can be attached to one host creates an artificial ceiling.

It also creates a brittle dependency: every enclosure, cable and hub converges on the same machine.

Physical concentration

Connecting every device to one host concentrates:

  • cabling;

  • power demand;

  • USB controllers;

  • enclosure management;

  • device events;

  • and physical maintenance.

The resulting system becomes increasingly difficult to modify safely.

Replacing one cable or enclosure may disturb unrelated devices. A host maintenance event affects the complete storage estate. Troubleshooting requires navigating a dense physical topology attached to one machine.

Distribution makes the physical design more modular.

Resource concentration

The central host would need to handle both:

  • raw device-management activity;

  • and the higher-level ZFS storage-control responsibilities.

That concentrates processor, memory and I/O pressure.

ZFS benefits substantially from memory and uses it deliberately for caching and storage operations. Adding every device-management workload to the same machine makes resource behaviour less predictable.

Distributing raw devices across back-end nodes moves some work away from the front-end control point.

Each node can handle its own:

  • device communication;

  • enclosure behaviour;

  • local operating-system activity;

  • maintenance;

  • and iSCSI target service.

HG000009 remains focused on ZFS and the presentation of controlled storage services.

What iSCSI provides

iSCSI carries SCSI block-storage commands over an IP network.

To the consuming host, an iSCSI logical unit appears as a block device. The consumer can place a filesystem, volume manager or another storage layer on top of it.

The main participants are:

  • the target, which provides the storage;

  • the initiator, which connects to and consumes it;

  • and the LUN, representing the logical block-storage unit presented by the target.

In my architecture:

  • the back-end nodes operate as iSCSI targets;

  • HG000009 operates as the initiator;

  • and ZFS manages the block devices presented through those sessions.

This creates a clear relationship between the node providing capacity and the node controlling the resulting storage service.

Horizontal expansion

A distributed iSCSI design supports horizontal growth.

Instead of expanding one host indefinitely, additional capacity can be introduced by adding:

  • another physical device;

  • another enclosure;

  • another back-end node;

  • or another iSCSI target.

This does not make scaling automatic.

The front-end system must still have sufficient network, processor and memory capacity. ZFS pool design must also account for how the new storage will be incorporated.

What changes is the physical constraint.

The architecture no longer depends on attaching every device to the same USB topology.

Capacity can grow through additional storage-providing nodes.

Resource isolation

Each back-end node creates a degree of resource isolation.

A device or enclosure performing maintenance can consume resources on its local node without placing the complete device-management burden on HG000009.

This separation can make system behaviour easier to understand.

If one back-end node becomes busy, the effect should be observable through:

  • the corresponding iSCSI session;

  • the associated network path;

  • and the ZFS devices dependent on that target.

The problem has a clearer boundary.

Without that separation, all device behaviour appears within one host, where contention between unrelated devices and services may be harder to isolate.

Independent maintenance

Distributed back-end nodes can undergo certain maintenance operations independently.

For example, one node may be:

  • restarting a target service;

  • replacing a device;

  • checking storage;

  • recovering from an enclosure problem;

  • or undergoing operating-system maintenance.

The effect depends on the redundancy and pool layout above it.

Distribution does not mean that a back-end node can disappear without consequence. If ZFS depends on a device presented by that node, loss of the iSCSI session appears as loss of that device.

The architecture must therefore be designed so that the resulting degraded state is understood and acceptable.

The advantage is not the elimination of failure.

It is the ability to isolate and reason about failure at the level of a particular back-end node or path.

Parallel operations

A centralised design can force some maintenance and recovery work through one resource bottleneck.

With distributed storage providers, different back-end nodes can perform operations at the same time.

That may include:

  • local device checks;

  • enclosure maintenance;

  • target-service work;

  • recovery activity;

  • and node-specific diagnostics.

Parallelism can reduce long serial maintenance paths.

It must still be controlled carefully. Running intensive operations simultaneously across many devices can saturate the shared network or place excessive load on HG000009.

Independent capability does not mean every operation should always happen concurrently.

It means the architecture is not inherently restricted to one physical host performing all work.

Explicit initiator-target relationships

iSCSI provides a defined relationship between the storage consumer and provider.

A target can restrict which initiators may access a LUN. The configuration can show:

  • which node exports the storage;

  • which initiator consumes it;

  • which LUN is presented;

  • and which session is active.

That visibility is valuable.

Raw storage should not be exposed to any system capable of reaching the network.

Each relationship should be intentional and documented.

Stable identifiers are also important. Storage should not depend on unpredictable device naming that changes after a reboot or reconnection.

The configuration must allow HG000009 to associate each presented device with the correct back-end target consistently.

Authentication

iSCSI can use CHAP authentication.

Depending on the configuration, this may authenticate the initiator to the target or support mutual authentication in which both ends authenticate.

Authentication helps prevent an unauthorised initiator from connecting merely because it can reach the target network.

It does not, by itself, provide encryption.

This distinction matters.

CHAP can establish that the connecting party possesses the expected secret. Standard iSCSI traffic may still be visible to an attacker with access to the network path.

The security model should therefore also rely on:

  • a dedicated or strongly segmented storage network;

  • narrow firewall rules;

  • controlled routing;

  • and, where required, a separate encryption mechanism.

Authentication is one control, not the complete security architecture.

A dedicated storage network

Storage traffic should not compete unnecessarily with ordinary client or administrative traffic.

A dedicated network or isolated segment provides several benefits:

  • predictable bandwidth;

  • clearer trust boundaries;

  • simpler access control;

  • easier monitoring;

  • and reduced exposure.

Only the required initiators and targets should participate.

The network should not become a general-purpose route between back-end systems.

If HG000009 has interfaces on several networks, IP forwarding should not accidentally turn it into a router between those segments.

Each interface should have a defined role.

The network becomes part of the storage system

Using iSCSI means that network reliability directly affects storage reliability.

A cable, interface, switch or configuration problem may appear to ZFS as a storage-device failure.

This is one of the most important trade-offs in the design.

With locally attached storage, the device path is physically short.

With iSCSI, the path includes:

  • the initiator;

  • the initiator’s network interface;

  • the network;

  • the target’s interface;

  • the target service;

  • the back-end operating system;

  • and the physical device.

Every component must be considered in the failure model.

This does not make network storage unsuitable.

It means the network can no longer be treated as an unrelated infrastructure service.

It is part of the storage architecture.

Latency and throughput

iSCSI introduces network latency and protocol overhead.

The practical effect depends on:

  • network speed;

  • switch behaviour;

  • interface quality;

  • frame size;

  • target performance;

  • workload;

  • queue depth;

  • and the characteristics of the physical devices.

A benchmark should not be limited to sequential throughput.

Storage workloads may involve:

  • small random reads;

  • synchronous writes;

  • metadata activity;

  • sustained transfers;

  • and mixed operations.

The system should be tested using workloads representative of the services it will provide.

A high headline throughput number does not guarantee predictable latency.

Multipathing and redundancy

Where availability requirements justify it, iSCSI can use multiple network paths.

Multipathing can protect against failure of:

  • an interface;

  • a cable;

  • a switch path;

  • or another network component.

However, multipathing adds complexity.

It requires:

  • correct path identification;

  • consistent target presentation;

  • appropriate failure detection;

  • tested path switching;

  • and careful interaction with the storage layer.

Adding a second path without testing failover can create the appearance of redundancy without reliable recovery behaviour.

In a home lab, multipathing is valuable as a learning exercise and may improve resilience. It should be introduced only when the complete failure path is understood.

Timeouts and failure detection

Storage-network failures require carefully considered timeout behaviour.

If failure is detected too quickly, a temporary network interruption may cause unnecessary device loss or recovery activity.

If detection takes too long, the storage stack may remain blocked while applications wait indefinitely.

Timeouts exist at several layers:

  • network;

  • iSCSI session;

  • SCSI device;

  • ZFS;

  • filesystem;

  • and application.

These layers must behave coherently.

The correct values depend on the environment and cannot safely be copied from a generic tuning guide without testing.

The design goal is predictable behaviour:

  • temporary interruptions are handled appropriately;

  • persistent failures become visible;

  • and recovery does not create ambiguity about the state of the data.

Write caching and data safety

Storage systems may cache writes at several points:

  • the application;

  • operating-system page cache;

  • ZFS;

  • the iSCSI target;

  • the back-end operating system;

  • the enclosure bridge;

  • and the physical device.

Each layer may report completion based on different assumptions.

If a component acknowledges a write before the data is safely stored, power loss or node failure can produce unexpected results.

The architecture must therefore understand:

  • which layers cache writes;

  • whether cache flushing is honoured;

  • which components have protected power;

  • and what durability guarantees the consumer actually receives.

Performance optimisation must not silently weaken the data-consistency model.

Start-up ordering

After a restart, HG000009 must re-establish its iSCSI sessions before ZFS can use the corresponding devices reliably.

This creates an explicit dependency chain:

Storage network available
        ↓
iSCSI sessions established
        ↓
Block devices visible
        ↓
ZFS pools imported
        ↓
Datasets and volumes available
        ↓
Client-facing services started

If client services start before the underlying storage is ready, the node may enter a partial or misleading state.

The operating system’s service-management configuration should encode these dependencies rather than rely on arbitrary delays.

A successful boot means more than reaching a login prompt.

It means the complete storage service has reached its defined healthy state.

Monitoring the complete path

Observability must cover both the storage and network layers.

Relevant information includes:

  • iSCSI session state;

  • reconnect events;

  • target availability;

  • network errors;

  • retransmissions;

  • throughput;

  • latency;

  • block-device errors;

  • ZFS pool health;

  • and back-end device behaviour.

When a ZFS device reports a problem, the investigation should be able to distinguish between:

  • a physical-disk issue;

  • an enclosure problem;

  • a target-service failure;

  • a back-end node failure;

  • and a network interruption.

Without that visibility, every problem may be described simply as “the storage disappeared.”

Operational ownership

Distribution creates clearer boundaries only if ownership is also clear.

The architecture should define responsibility for:

  • physical devices;

  • enclosures;

  • back-end operating systems;

  • target configuration;

  • the storage network;

  • initiator configuration;

  • ZFS pools;

  • and client-facing services.

In a home lab, one person may perform every role.

The separation is still valuable because it creates a model that can scale conceptually to larger environments.

In production, those responsibilities may belong to different teams. Explicit boundaries reduce the risk that each group assumes another owns the problem.

The trade-off

iSCSI does not make the design simpler in every respect.

It adds:

  • network dependency;

  • initiator and target configuration;

  • authentication;

  • service ordering;

  • timeout behaviour;

  • and additional observability requirements.

The choice is justified only if those costs are outweighed by the benefits:

  • horizontal expansion;

  • resource distribution;

  • physical modularity;

  • clearer back-end boundaries;

  • independent maintenance;

  • and more flexible storage placement.

For this architecture, I considered that trade worthwhile.

A deliberate architectural boundary

The central reason for using iSCSI is separation.

The back-end nodes provide raw block-storage resources.

HG000009 provides the storage control plane through ZFS.

Consumers receive controlled storage services from the front-end layer.

Each part has a clear responsibility.

This does not remove complexity. It organises complexity into boundaries that can be documented, secured, monitored and tested.

That is the purpose of the design.

iSCSI is not merely the protocol carrying blocks across a network.

It is the deliberate boundary between raw storage capacity and the system responsible for turning that capacity into an understandable and recoverable service.

AI Assistance Disclosure

This article is based on my original ideas, experience, analysis and conclusions. Artificial intelligence tools were subsequently used as editorial and research assistants to review grammar and wording, improve structure and presentation, organise some arguments into clearer logical sections, and help review references to legal, regulatory and technical concepts.

Where relevant, factual and regulatory references were checked against the sources cited in the article. AI assistance does not replace professional legal, regulatory, financial or technical advice, and the final selection, interpretation, opinions and conclusions presented here remain my own.

Comments

Popular posts from this blog

Movies - The Bubble (2022)

  Back to Evolution (2001) .

IT - Fixing Windows Error 1327: Account Restrictions Are Preventing This User from Signing In

Fixing Windows Error 1327: Account Restrictions Are Preventing This User from Signing In Introduction Error 1327, “Account restrictions are preventing this user from signing in,” is a perplexing and disruptive issue that occurs on some Windows 10 and Windows 11 machines. The message typically appears at login or while connecting to remote resources, like shared folders, network drives, or remote desktops. Table of Contents Symptoms of Error 1327 Common Causes Step-by-Step Troubleshooting Advanced Fixes Automation via PowerShell Prevention Tips Further Reading Symptoms of Error 1327 Users experiencing this error may encounter one or more of the following: Login screen fails after credentials are entered. Error message appears when accessing mapped drives or network resources. Remote Desktop Connection (RDP) is rejected with the 1327 message. Group Policy logon restrictions silently block access. Co...

ATA Drive Capacity Limitations

ATA interface versions up through ATA-5 suffered from a drive capacity limitation of about 137GB (billion bytes). Depending on the BIOS used, you can further reduce this limitation to 8.4GB, or even as low as 528MB (million bytes). This is due to limitations in both the BIOS and the ATA interface, which when combined create even further limitations. To understand these limits, you have to look at the BIOS (software) and ATA (hardware) interfaces together. NOTE In addition to the BIOS/ATA limitations discussed in this section, various operating system limitations exist. These are described later in this chapter. The limitations when dealing with ATA drives are those of the ATA interface as well as the BIOS interface used to talk to the drive. A summary of the limitations is shown in Table 7.12. Table 7.12. ATA/IDE Capacity Limitations for Various Sector Addressing Methods Sector Addressing Method Total Sectors Calculation Maximum Total Sectors Maximum Capacity (Byte...