Based on a LinkedIn post originally published on 6 January 2026
Over several months, I designed my home lab as an end-to-end platform and storage architecture exercise.
The objective was not to assemble a collection of fashionable tools or run as many services as possible at home. I wanted to create a coherent system in which decisions about hardware, storage, networking, security, reliability and operations could be examined together.
This distinction matters.
A collection of independently installed technologies may demonstrate familiarity with individual products. It does not necessarily demonstrate how those products interact, where the trust boundaries are, how failures propagate or whether the resulting platform can be operated predictably.
I wanted the home lab to answer those broader questions.
A complete system rather than a collection of tools
The architecture covers several interconnected areas:
ZFS-backed storage;
iSCSI-based shared storage;
network segmentation;
clustering and failover concepts;
platform and control-plane services;
observability;
operational access;
and security controls embedded throughout the design.
Each part must be understood in relation to the others.
Storage decisions affect network traffic. Network design influences security boundaries. Security controls affect administration and recovery. Hardware limitations shape availability and performance. Observability must cover the complete system rather than merely collect isolated metrics from individual components.
This is where a home lab becomes an architecture exercise rather than an installation exercise.
Deliberately modest hardware
The hardware foundation consists primarily of older Apple Mac minis.
This was a deliberate constraint.
It would have been possible to purchase newer and more powerful equipment, but abundant resources can conceal architectural weaknesses. When processor capacity, memory, storage connectivity and network interfaces are limited, design decisions become more visible.
You must consider questions such as:
Which services genuinely need dedicated resources?
Which workloads can safely share a host?
Where could contention occur?
What happens when one node becomes unavailable?
Which component represents a single point of failure?
How much redundancy is justified?
What level of performance is realistically achievable?
Which trade-offs are acceptable within the available capacity?
Working within constraints forces discipline. It requires realistic capacity planning and makes it more difficult to substitute hardware for sound design.
This closely resembles real production environments. Commercial platforms rarely have unlimited budgets or unlimited resources. Engineering is largely the process of making responsible decisions within constraints.
ZFS as the storage foundation
ZFS provides the principal storage layer.
The choice was based on capabilities such as data integrity, snapshots, replication and predictable storage management. It also makes important architectural concerns visible: memory allocation, caching behaviour, disk organisation, failure handling and recovery planning.
Storage cannot be treated as an opaque box.
The design must account for what happens when a disk fails, how corrupted data is detected, how capacity is expanded and how services continue operating during maintenance. Snapshot and replication strategies must also reflect the actual recovery requirements rather than exist merely because the technology supports them.
Using ZFS encourages explicit thinking about the lifecycle of data:
Where is it created?
Where is it stored?
How is its integrity protected?
How is it presented to consuming systems?
How is it replicated?
How is it recovered after a failure?
When and how is it eventually removed?
These are production questions, regardless of whether the platform is running in a data centre, a public cloud or a home lab.
Separating storage through iSCSI
The architecture uses iSCSI to present shared block storage across the network.
This separates the systems consuming storage from the physical devices providing it. It also makes the network an explicit part of the storage architecture.
That creates useful engineering challenges.
Storage traffic must be isolated appropriately. Authentication and access must be controlled. Latency and throughput must be understood. Failure behaviour must be tested instead of assumed. The system must also make it clear which component owns a device and what happens when connectivity is interrupted.
Using iSCSI is therefore not simply a matter of exporting a disk from one machine and mounting it on another. It requires consideration of ownership, consistency, authentication, network reliability and recovery.
Designing failure domains
One of the central objectives was to define clear failure domains.
A reliable platform is not one in which nothing ever fails. It is one in which failures are anticipated, contained, observable and recoverable.
For every significant component, the architecture should answer:
What depends on it?
What happens when it stops responding?
Can another component continue the service?
Is the failure immediately visible?
Does recovery happen automatically or require intervention?
Could an attempted recovery cause data corruption or another failure?
Is there a documented procedure for restoring normal operation?
Clustering and failover are useful only when they address a clearly understood failure scenario. Adding redundant components without understanding their dependencies can produce the appearance of resilience while preserving the same underlying single point of failure.
The lab therefore treats availability as a system property rather than a product feature.
Network segmentation and trust boundaries
The network design separates traffic according to purpose.
Storage, administration, platform services and client access do not automatically belong on the same network. Each has different performance requirements and a different security profile.
Segmentation makes the intended relationships explicit:
which systems may communicate;
which services may be reached;
which paths are administrative;
which paths carry storage traffic;
and where controls should be enforced.
This reduces accidental exposure and makes troubleshooting easier. When every interface and network has a defined role, unexpected traffic becomes easier to identify.
Trust is not granted merely because systems are physically located in the same room. Access is permitted because it is required and intentionally designed.
Security as an architectural property
Security controls were included from the beginning rather than added after implementation.
This means considering identity, authentication, network exposure, administrative access, data protection, logging and recovery while the architecture is still being designed.
Retrofitting security often creates awkward exceptions and inconsistent controls because important assumptions have already become embedded in the platform.
Designing security from the start allows more fundamental questions to be addressed:
Which services need to exist?
Who or what should access them?
From which network locations?
Using which identities and credentials?
What evidence should be recorded?
How will access be revoked?
What happens if a component is compromised?
Security becomes more manageable when the architecture begins with a narrow and explicit trust model.
Observability with a purpose
The lab also includes observability and platform-control services, but the intention is not to collect every available metric.
Useful observability begins with operational questions.
What information is required to determine whether the platform is healthy? Which conditions require intervention? Which measurements help diagnose a failure? What historical information is needed to understand capacity or performance trends?
Metrics, logs and alerts should support decisions. Data that is never examined or connected to an operational outcome creates cost and noise without improving reliability.
The observability design therefore follows the architecture’s failure domains and service boundaries. It should make the state of the system understandable, not merely produce dashboards.
Documenting decisions and trade-offs
I approach the lab in the same way I would approach a production platform:
dependencies are explicit;
ownership is defined;
data, control and failure domains are separated;
security and reliability are first-class concerns;
and important decisions are documented and reviewed.
Documentation is not simply a final diagram showing what was built. It should explain why the design exists, which alternatives were considered and what trade-offs were accepted.
That context becomes essential when the architecture changes or when a failure challenges the original assumptions.
The real objective
This project is not about “running everything at home.”
It is an opportunity to apply systems thinking across the complete technology stack—from physical hardware and storage to networking, platform services, operations and security.
The value lies in the connections between these areas.
A storage decision is also a networking decision. A network decision is also a security decision. A clustering decision is also an operational and data-consistency decision. An observability decision should reflect all of them.
By building the platform incrementally, testing its assumptions and documenting what works and what does not, the home lab becomes a practical environment for exploring the realities of platform engineering.
That is the purpose of the project: not to reproduce a commercial data centre in miniature, but to practise disciplined system design under realistic constraints.
This article is based on my original ideas, experience, analysis and conclusions. Artificial intelligence tools were subsequently used as editorial and research assistants to review grammar and wording, improve structure and presentation, organise some arguments into clearer logical sections, and help review references to legal, regulatory and technical concepts.
Where relevant, factual and regulatory references were checked against the sources cited in the article. AI assistance does not replace professional legal, regulatory, financial or technical advice, and the final selection, interpretation, opinions and conclusions presented here remain my own.
Comments
Post a Comment