Based on a LinkedIn post originally published on 6 March 2026
In the early 1990s, long before network connectivity became an ordinary assumption, I worked with Xenix systems in an environment where conventional local-area network infrastructure either did not exist or was not economically accessible.
We still needed the machines to exchange information.
So we built a network using the resources available to us: serial cables, Xenix servers and UUCP.
The result would appear primitive when compared with a modern Ethernet network, but it provided the capabilities the environment required:
- files could be transferred between systems;
- messages could be exchanged;
- processing jobs could be submitted to other machines;
- and information could propagate across several independently operating servers.
In effect, we created a small distributed system before that terminology became part of everyday technology discussions.
The system did not depend on permanent connectivity. It depended on every machine understanding what should happen when a connection eventually became available.
A network made from serial links
Each server was connected to one or more other systems through serial communication. UUCP—originally named for UNIX-to-UNIX Copy—provided the mechanism for queuing and transferring work between them.
The topology could be simple and still support communication beyond one direct neighbour:
+------------+ serial +------------+
| Xenix |<------------------>| Xenix |
| Server A | | Server B |
+------------+ +------------+
|
| serial
|
+------------+
| Xenix |
| Server C |
+------------+
|
| serial
|
+------------+
| Xenix |
| Server D |
+------------+A file did not necessarily travel directly from its origin to its final destination. It could move from one system to another as links became available, following a defined route through the environment.
Each machine could continue operating independently while transfers waited in local queues.
Store now, forward later
UUCP used a store-and-forward model.
When a system needed to send a file, message or job, it placed the work in a queue. The transfer occurred when the remote machine could be contacted. If the destination was not the directly connected neighbour, the data could be forwarded through intermediate systems.
Create transfer request
|
v
Place work in local queue
|
v
Attempt connection
|
+---- Link unavailable ----> Wait and retry
|
v
Transfer to next system
|
+---- Final destination? -- No --> Queue for forwarding
|
Yes
|
v
Deliver and record completionThe absence of an immediate connection was not an exceptional event. It was a normal state for which the system had been designed.
This created a fundamentally different expectation from modern interactive networking. A successful request did not mean that the remote system had already received the information. It meant that the local system had accepted responsibility for attempting the transfer.
No illusion of real-time communication
There was no assumption of instant delivery.
A serial link might be occupied, slow or temporarily unavailable. A remote server might not respond. An intermediate machine could receive work but be unable to forward it immediately.
Everyone working with the system understood that delay was part of the model.
This honesty had an important architectural consequence: the applications and operational processes could not silently depend on permanent connectivity.
A system that admits communication may be delayed can define what delayed delivery means. A system that pretends connectivity is permanent often discovers its real assumptions only during an outage.
The distinction between accepted, queued, transferred, forwarded and delivered was visible. Those states mattered because an operator needed to know whether intervention was required.
Failure was part of normal operation
UUCP forced us to think explicitly about unreliable links.
A failed connection did not necessarily mean the operation had failed permanently. It might mean only that the work needed to remain queued until the next attempt.
Reliability emerged from several simple mechanisms:
- Queues preserved work while the next system was unavailable.
- Retries allowed transient communication failures to recover without recreating the request manually.
- Logs recorded what had been attempted and what had happened.
- Explicit routing made the expected path understandable.
- Independent systems continued operating while disconnected.
The design did not eliminate failure. It prevented ordinary failure from automatically becoming data loss or an operational crisis.
Link available:
queue -> transfer -> confirmation
Link unavailable:
queue -> retain state -> log attempt -> retry later
Permanent problem:
repeated failure -> visible evidence -> operator actionAn early lesson in eventual consistency
At any given moment, the systems did not necessarily contain identical information.
A file might exist on Server A while its transfer to Server B remained queued. Server B could later receive it while the copy intended for Server C still awaited forwarding.
The environment was temporarily inconsistent, but it had a defined process through which the required state would propagate.
In modern terminology, this resembles eventual consistency:
Time T1 Server A: new data Server B: previous data Server C: previous data Time T2 Server A: new data Server B: new data Server C: previous data Time T3 Server A: new data Server B: new data Server C: new data The systems converge as communication succeeds.
The important property was not that every node was synchronised continuously. It was that transfers converged safely, predictably and visibly.
That final word matters. Convergence without sufficient observability becomes hope. We needed to know whether work remained queued, had failed repeatedly or had reached its destination.
Modern patterns in an old environment
Looking back, many concepts associated with contemporary distributed architecture were already present:
- Asynchronous communication: the sender did not wait for the complete end-to-end operation.
- Store-and-forward messaging: intermediate systems retained data until forwarding became possible.
- Producer-consumer decoupling: the originating system could create work without the destination being continuously available.
- Durable queues: pending transfers survived temporary communication failures.
- Retry behaviour: transient failures were expected and managed.
- Distributed state: different systems temporarily held different versions of reality.
- Operational observability: logs and queue state were essential for determining whether the system was progressing.
Today, we implement related ideas using message brokers, event streams, replication pipelines, distributed logs and cloud services. The bandwidth is greater, the interfaces are richer and much of the complexity is hidden behind abstraction.
The underlying problem remains familiar: independent systems must exchange work even though communication can be delayed or fail.
Abstraction changes visibility
Modern platforms allow teams to build distributed systems without implementing every low-level communication mechanism themselves. That is an enormous advantage.
But abstraction can create an illusion that the underlying failure modes have disappeared.
A managed message service still has queues. Messages can still be delayed, duplicated, rejected or processed out of order. Consumers can fail after performing work but before recording completion. Network partitions still exist, even when nobody sees a serial cable.
Abstraction removes implementation work. It does not remove the need to understand delivery semantics, failure states and recovery behaviour.
The constraints of the Xenix and UUCP environment made those realities impossible to ignore. Every delay was visible. Every queued transfer had a physical and operational meaning. Every unreliable link had to be accommodated deliberately.
Reliability begins with honest assumptions
A reliable design does not begin by assuming that every dependency will remain available.
It begins by asking:
- What happens when the destination cannot be reached?
- Where is unfinished work retained?
- How is a retry scheduled?
- Can repeated processing cause damage?
- How does an operator see that progress has stopped?
- What evidence confirms successful delivery?
- How long can the systems remain inconsistent?
These questions applied to a few Xenix servers connected by serial cables. They still apply to services distributed across regions and cloud providers.
The technologies have changed. The need to design for imperfect communication has not.
Reliability is not created by pretending failure is exceptional. It is created by making failure an explicit, observable and recoverable part of normal system behaviour.
This article is based on my original ideas, experience, analysis and conclusions. Artificial intelligence tools were subsequently used as editorial and research assistants to review grammar and wording, improve structure and presentation, organise some arguments into clearer logical sections, and help review references to legal, regulatory and technical concepts.
Where relevant, factual and regulatory references were checked against the sources cited in the article. AI assistance does not replace professional legal, regulatory, financial or technical advice, and the final selection, interpretation, opinions and conclusions presented here remain my own.
Comments
Post a Comment