Based on a LinkedIn post originally published on 30 January 2026
While working in New Zealand, I supported a web-facing portal designed with strong redundancy.
The environment included:
local high availability;
multiple portal nodes;
production and disaster-recovery capacity;
and remote replication between the primary and DR environments.
From an operational perspective, it appeared to be an excellent candidate for rolling maintenance.
Individual servers could be removed from service, patched and returned while traffic continued through the remaining nodes.
As part of the regular security-maintenance cycle, we scheduled the deployment of a Solaris 10 patch.
The process followed the expected progression:
Laboratory environment.
Development and test.
Production.
The patch behaved correctly in every non-production environment.
We therefore proceeded to production.
That was when the supposedly routine change stopped being routine.
Starting with the lowest-risk node
I began with the first server in the disaster-recovery environment.
This was deliberate.
The DR node provided a controlled first exposure to a production configuration while reducing the risk to the live primary service. If the patch behaved as expected, we could proceed one node at a time.
The deployment began normally.
Then the server stopped progressing.
It did not complete the patch installation. It did not return a useful failure. It did not recover gracefully.
The system was effectively wedged.
A failure around the SSH service
Initial investigation pointed towards the SSH service.
On Solaris 10, services were managed through the Service Management Facility, commonly known as SMF. Instead of relying solely on traditional start-up scripts, SMF represented services and their dependencies through a managed framework.
In principle, this provided several advantages:
consistent service state;
dependency management;
automatic restart behaviour;
centralised administration;
and clearer diagnostic information.
During this patch deployment, however, the interaction involving the SSH service left the server in an unusable state.
The behaviour made little sense.
The same patch had already worked in other environments using the same operating-system release and apparently equivalent configuration.
With the patch process no longer progressing and the system unable to recover normally, the only practical option was to restart the server.
That was already uncomfortable.
The more important question was whether the same failure would occur on every remaining portal node.
Stop after the first unexpected result
One of the most important maintenance disciplines is knowing when to stop.
A change plan may authorise patching ten or twenty systems. That does not mean the engineer should continue mechanically after the first unexplained failure.
The first DR node had provided new evidence:
the tested procedure was not universally safe;
production contained a condition absent from the earlier environments;
the server could hang;
and a restart might be required.
Continuing immediately would have converted one contained failure into a potential service-wide incident.
I paused the rollout and contacted Sun for an explanation.
The vendor’s explanation
After discussion and investigation, the explanation emerged.
There was a known issue involving the Solaris SSH service that could be triggered by particular deployment patterns—including the one used in our environment.
The reliable workaround was to apply the patch and restart the server.
That led to the obvious question:
Why was this requirement not included in the patch documentation?
The answer was uncomfortable.
The problem had been discovered as a side effect of deploying the patch itself.
The update corrected one issue but exposed another interaction in certain real-world configurations. Those systems required a restart, even though the patch had not initially been presented that way.
The behaviour did not appear everywhere. That was why our laboratory, development and test environments had completed successfully.
Production had found the condition that the earlier stages did not reproduce.
Test environments reduce risk; they do not eliminate it
A successful laboratory test is valuable.
A successful development and test deployment is even more valuable.
Neither proves that production will behave identically.
Production environments differ in ways that are not always obvious:
workload;
uptime;
historical configuration;
service state;
device and driver combinations;
installed software;
dependency timing;
network relationships;
accumulated patches;
operational data;
and interaction with real traffic.
Two servers may report the same operating-system version and still contain meaningful differences.
The production system has also experienced a history that the test environment may not reproduce. Services may have been running for months. Resources may be in states that never occur during short testing cycles.
The patch was not tested incorrectly.
The available tests simply did not contain the condition that triggered the problem.
Configuration equivalence is difficult
Organisations often describe a test environment as “the same as production.”
That statement usually means “similar enough for the intended purpose.”
Perfect equivalence is difficult and often prohibitively expensive.
A non-production environment may differ in:
scale;
data volume;
traffic;
external integrations;
hardware;
redundancy;
security controls;
or maintenance history.
Even when configuration management applies the same declared settings, runtime state can diverge.
The important practice is not to claim that environments are identical. It is to document known differences and understand which risks those differences leave untested.
The value of the production canary
The first DR node effectively served as a production canary.
A canary is a deliberately limited first deployment used to expose the change to realistic conditions before wider rollout.
For patching, a suitable canary should:
represent the production configuration;
have limited impact if it fails;
be observable;
support recovery;
and provide enough time for unexpected behaviour to become visible.
The canary does not guarantee that every later node will behave identically.
It provides another checkpoint between testing and broad production deployment.
In this case, it performed exactly that function. The failure occurred on one controlled node instead of across the complete portal estate.
Changing the maintenance plan
Once the vendor confirmed that restarts were required, the rollout plan had to change.
The original expectation was that each node could be patched and returned without an exceptional restart requirement.
The revised process became:
Remove one node from service.
Confirm that traffic had shifted to the remaining instances.
Apply the patch.
Restart the server.
Validate the operating system and required services.
Return the node to the pool.
Observe service behaviour.
Proceed to the next node only after the previous one was healthy.
Every portal server needed to be restarted.
Nobody was pleased.
A restart introduces additional time and risk. It can expose unrelated start-up problems, dependency issues or hardware faults that remained hidden while the server was continuously operating.
Nevertheless, the redundancy design allowed the work to continue without taking the portal offline.
Redundancy provided operational freedom
High availability is often discussed primarily as protection against unplanned failure.
It is equally valuable during planned maintenance.
Because traffic could shift between nodes, each server could be removed, restarted and validated separately.
The architecture provided room to respond to new information.
Without redundancy, the unexpected restart requirement might have forced a choice between:
accepting an outage;
postponing an important security patch;
or continuing under significant uncertainty.
Redundancy did not prevent the patch problem.
It contained the consequence.
That is an important distinction.
A restart is not a complete recovery plan
“Reboot the server” is sometimes treated as a trivial recovery action.
In production, it requires preparation.
Before restarting a node, the engineer should understand:
whether traffic has been drained;
whether active sessions will be affected;
whether data has been flushed safely;
whether dependent services can tolerate the loss;
whether the server is expected to boot automatically;
whether storage and network dependencies will return;
whether applications start in the correct order;
and how the node will be validated afterward.
A server that responds to a network check after boot is not necessarily ready for production traffic.
Validation should include the complete service path.
Patch notes are necessary but incomplete
Patch documentation is an essential source of information.
It may describe:
corrected defects;
affected components;
prerequisites;
known issues;
installation instructions;
restart requirements;
and rollback procedures.
However, release notes can only document behaviour known and understood when they are produced.
A rare interaction may be discovered only after customers deploy the update across a wider variety of environments.
This is not an argument for ignoring documentation.
It is an argument for recognising its limits.
The change plan must account for the possibility that the documentation is incomplete.
Vendor escalation is part of engineering
The decision to contact Sun before proceeding was not an admission that the internal team lacked competence.
Vendor escalation is a legitimate part of managing enterprise technology.
The vendor may have access to:
internal defect records;
source-level knowledge;
reports from other customers;
engineering specialists;
and information not yet included in public documentation.
The internal team contributes the environment-specific evidence. The vendor contributes product-specific knowledge.
A useful escalation should provide:
the exact patch and operating-system level;
the observed behaviour;
relevant logs and service state;
the deployment sequence;
differences between successful and failed systems;
and the business impact.
The quality of that evidence influences how quickly the issue can be understood.
Security urgency must be balanced with operational risk
Security patches address known vulnerabilities or reduce exposure.
Delaying them indefinitely creates risk.
Applying them without sufficient control can create another kind of risk: loss of service, data or recoverability.
Patch management therefore requires judgement.
The decision should consider:
the severity and exploitability of the vulnerability;
current exposure;
available compensating controls;
confidence in the patch;
recovery readiness;
redundancy;
maintenance timing;
and the consequence of failure.
There is no universal rule that every patch must be applied immediately or that operational caution should always delay deployment.
The responsible decision depends on evidence and context.
A stronger patching process
The incident reinforced several practices that make production patching safer.
Maintain an accurate inventory
Know which systems exist, what they run and which services depend on them.
Understand environment differences
Document where production differs from laboratory and test systems.
Use staged deployment
Move progressively from low-risk environments to representative production nodes.
Establish stop conditions
Define which unexpected results require the rollout to pause.
Prepare recovery
Verify backups, rollback options, console access and restart procedures before beginning.
Observe the full service
Monitor user-facing behaviour, not only whether the patch command completed.
Patch one failure domain at a time
Do not remove redundant capacity simultaneously.
Contact the vendor when evidence contradicts expectations
Do not continue merely because the change window is already open.
Record what happened
Update procedures and known-issue records so the next maintenance cycle benefits from the experience.
The difference between procedure and judgement
The original process followed the expected procedure:
test the patch;
progress through environments;
begin with a lower-risk node;
and patch systems individually.
The procedure was sound.
It did not prevent an unknown product interaction.
What protected the service was judgement:
starting with the DR node;
recognising that the hang was significant;
stopping the rollout;
escalating to the vendor;
changing the plan;
and monitoring traffic while restarting nodes one at a time.
Procedures create a safe baseline.
Judgement determines what to do when reality no longer matches the procedure.
Production always has another edge case
The lesson was not that laboratory and test environments are useless.
They significantly reduce risk and catch many problems before production.
The lesson was that they cannot prove the absence of every failure.
Production contains real scale, real history, real integrations and real operational state. It will eventually expose a combination that earlier testing did not include.
Good engineering accepts that possibility.
The goal is not to design a process in which surprises become impossible.
The goal is to ensure that a surprise:
appears first in a limited scope;
becomes visible quickly;
does not propagate uncontrollably;
and can be recovered without losing the service.
Patching is not merely the execution of a technical command.
It is a risk-management exercise.
Sometimes the most important skill is not avoiding the unexpected.
It is containing it when it arrives.
This article is based on my original ideas, experience, analysis and conclusions. Artificial intelligence tools were subsequently used as editorial and research assistants to review grammar and wording, improve structure and presentation, organise some arguments into clearer logical sections, and help review references to legal, regulatory and technical concepts.
Where relevant, factual and regulatory references were checked against the sources cited in the article. AI assistance does not replace professional legal, regulatory, financial or technical advice, and the final selection, interpretation, opinions and conclusions presented here remain my own.
Comments
Post a Comment