Based on a LinkedIn post originally published on 9 March 2026
While working for a bank in New Zealand, we operated a fleet of mid-range servers running SunOS and Solaris, ranging from SunOS 5 through Solaris 9.
These systems supported important workloads, so their boot architecture, storage resilience and recoverability were operational concerns rather than implementation details.
One of the strengths of the Sun environment was its flexible boot and disk architecture. When configured correctly, a server could boot from different storage devices without depending permanently on one specific internal disk.
This flexibility eventually led us to ask a deceptively simple question:
Once reliable SAN storage was available, why should the servers retain local boot disks at all?
The original local-disk model
Before the SAN migration, each server contained its own internal disks. A basic boot structure could exist on each disk, while the remaining disk capacity was protected through SunOS software mirroring.
The operating system itself could run from mirrored volumes, which reduced dependence on one individual disk.
+--------------------- Sun server ---------------------+ | | | Internal disk A Internal disk B | | +-------------+ +-------------+ | | | Boot layout | | Boot layout | | | | OS mirror A |<------------>| OS mirror B | | | +-------------+ +-------------+ | | | +------------------------------------------------------+
This arrangement provided redundancy, but every server still contained mechanical boot devices that required monitoring, replacement and model-specific maintenance.
A failed internal disk might not stop the operating system immediately if the mirror remained healthy, but it still created a degraded condition and required physical intervention.
The arrival of SAN storage changed the question
When the bank introduced a proper storage-area network, the infrastructure gained centrally managed storage with RAID protection, redundant connectivity and enterprise monitoring.
Initially, it would have been natural to use the SAN only for application data while leaving the operating systems on local disks.
That would have preserved a familiar separation:
Local disks -> Operating system and boot SAN storage -> Application and business data
But the distinction was no longer necessarily beneficial.
If the SAN could provide more resilient and more observable storage than the internal disks, keeping local boot devices meant preserving a separate class of hardware and failure mode inside every server.
The potential simplification was significant:
+---------------+ redundant FC paths +---------------+ | Sun server |==============================| SAN storage | | | | | | No local | RAID-1 boot LUN | Centralised | | boot disks |<-----------------------------| protection | +---------------+ +---------------+
Validate the idea on one system
Removing internal boot disks from an entire server fleet is not the sort of change that should begin with broad confidence and a mass rollout.
We started with one server and treated the migration as a sequence of independently verifiable steps.
The initial objective was not merely to demonstrate that Solaris could see a SAN LUN. We needed to establish that the complete system could:
- access the boot device through redundant Fibre Channel paths;
- mirror the operating system safely during migration;
- boot through the system firmware from SAN-attached storage;
- run its normal workload reliably;
- and restart successfully after the internal disks had been physically removed.
A successful first reboot was necessary, but not sufficient. The migration was complete only when the server could operate normally and recover using the SAN as its sole boot-storage path.
The validated migration process
The procedure developed through the first migration contained nine main steps.
1. Provide redundant Fibre Channel connectivity
We installed either two single-port Fibre Channel cards or one dual-port card. The purpose was to avoid replacing a local-disk dependency with a single storage-network dependency.
The host needed multiple Fibre Channel paths so that the loss of an individual port, cable or fabric path would not automatically remove access to the boot device.
2. Create and present a protected boot LUN
We created a RAID-1 LUN on the storage system and presented it to the selected host.
The mapping was deliberate and host-specific. The server received the storage capacity required for its operating system without gaining indiscriminate access to unrelated devices.
3. Extend the SunOS software mirror
We created a two-way or three-way SunOS software mirror that initially included both the existing internal storage and the new SAN-attached device.
This produced a controlled transition state:
Migration mirror
+-----------------+
| Internal disk A |
+-----------------+
|
+---------- SunOS software mirror
|
+-----------------+
| Internal disk B |
+-----------------+
|
+---------- Synchronisation
|
+-----------------+
| SAN boot LUN |
+-----------------+The operating system contents could be synchronised to the SAN while the existing boot devices remained available.
4. Remove the internal disks from the mirror
Once synchronisation and validation were complete, we removed the internal disks from the SunOS mirror logically.
At this stage, the hardware remained physically present, but the SAN LUN became the operating system’s active storage device.
5. Reconfigure firmware and operating-system boot settings
The system firmware and Solaris configuration were updated so that the machine would locate its boot device through the Fibre Channel infrastructure.
This step connected several architectural layers:
System firmware
|
Fibre Channel adapter
|
SAN path and multipath configuration
|
Presented RAID-1 LUN
|
Solaris boot environment6. Reboot from the SAN
The first SAN-based reboot demonstrated that the firmware, host-bus adapter, storage paths, LUN presentation and operating-system configuration formed a complete boot chain.
It worked.
7. Run the normal workload
We then operated the server under its ordinary workload for several days.
This observation period allowed us to look beyond basic boot success and examine stability, storage-path behaviour and application operation under realistic conditions.
8. Shut down and remove the internal disks
After the server had demonstrated stable operation, we shut it down and physically removed the internal disks.
This was the point at which the simplification became real. The server no longer retained its former storage devices as an undeclared fallback.
9. Perform the decisive reboot
Finally, we restarted the machine using only the SAN-attached boot storage.
The server booted and returned to operation successfully.
The final test removed the original escape route. Success meant the system could start, operate and recover with no dependency on internal boot disks.
From one validated server to the entire fleet
Once the complete process had been demonstrated, we migrated the remaining Sun servers progressively.
The sequence mattered. We did not treat the first technical success as permission to change every machine simultaneously. The process became repeatable because its dependencies, checks and recovery considerations had been identified on the initial system.
Single-server proof
|
v
Operational observation
|
v
Documented migration procedure
|
v
Progressive fleet rollout
|
v
Standard SAN-boot architectureThe operational impact
The improvement was immediate and measurable.
Improved reliability and availability
The boot devices now benefited from RAID protection, redundant Fibre Channel connectivity and the operational capabilities of the central storage platform.
Removal of local-disk failure modes
Internal boot disks could no longer fail because they were no longer present. The servers had fewer mechanical components and fewer degraded local-mirror conditions requiring physical attention.
Reduced server hardware complexity
The systems became simpler internally. Boot-storage resilience moved into the architecture already designed to provide resilient storage.
Simpler maintenance and replacement
Hardware maintenance no longer had to preserve or rebuild local operating-system disks in the same way. Storage identity and server hardware became less tightly coupled.
Centralised monitoring
The boot devices became visible through the same storage-management and monitoring capabilities used for other SAN resources.
Redundancy should remove dependencies, not multiply them
It is easy to respond to reliability requirements by adding another component.
More components can improve resilience when they eliminate a single point of failure. They can also create more paths, more states and more operational knowledge that must be maintained.
In this case, the SAN already provided a stronger storage foundation. Retaining local disks would have maintained two separate storage models:
WITH LOCAL BOOT DISKS Per-server disk hardware + local software mirroring + local failure replacement + SAN infrastructure + SAN monitoring SAN-ONLY BOOT Redundant Fibre Channel paths + centrally protected boot LUN + central storage monitoring
The simpler model did not reduce protection. It concentrated protection in the layer best equipped to provide it.
Use native capability before adding another layer
The design succeeded because it combined capabilities that already existed:
- Sun systems could boot from properly configured external storage;
- SunOS software mirroring supported a controlled migration path;
- the firmware could be configured to locate the SAN boot device;
- Fibre Channel provided redundant storage connectivity;
- and the storage system provided RAID protection and central management.
No additional orchestration product was needed to compensate for an architecture that the existing platform could already support.
The engineering work lay in understanding those capabilities, combining them safely and validating the entire boot chain under real operating conditions.
Good infrastructure design is not always about what else can be added. Sometimes the most effective improvement begins by asking what can be removed safely.
This article is based on my original ideas, experience, analysis and conclusions. Artificial intelligence tools were subsequently used as editorial and research assistants to review grammar and wording, improve structure and presentation, organise some arguments into clearer logical sections, and help review references to legal, regulatory and technical concepts.
Where relevant, factual and regulatory references were checked against the sources cited in the article. AI assistance does not replace professional legal, regulatory, financial or technical advice, and the final selection, interpretation, opinions and conclusions presented here remain my own.
Comments
Post a Comment