Based on a LinkedIn post originally published on 6 February 2026
A few years ago, I joined a very large start-up to manage a DevOps team.
The company’s product was extremely popular and clearly successful. The platform had grown rapidly to support increasing demand across several regions.
As a newcomer, I asked what I thought was a straightforward question:
Where is the architectural view of this environment?
There was no clear answer.
I then asked where I could find documentation explaining what existed, how the components interacted and how the platform was actually deployed.
Again, there was no authoritative answer.
Everyone understood a part
It gradually became apparent that no one had a complete view of the system.
Individual engineers understood their own areas. Application developers knew their services. Platform specialists understood particular infrastructure components. SREs knew the systems involved in the incidents they handled. Database specialists understood the data services they supported.
Each view was locally valuable but globally incomplete.
The platform had grown organically. Components had been introduced to solve immediate problems, teams had created their own operating practices and dependencies had accumulated over time.
There was no explicit, shared model showing how everything fitted together.
Application team Platform team SRE team
| | |
v v v
Local service view Infrastructure view Incident view
\ | /
\ | /
+-------- No shared system model ----+This situation is more common than many organisations would like to admit.
Reconstructing the system
I began interviewing everyone I could:
- software engineers;
- SREs;
- platform specialists;
- database engineers;
- infrastructure engineers;
- and people responsible for operational support.
The purpose was not to produce a visually impressive diagram.
I needed to understand what actually existed.
For each area, I tried to establish:
- which components were deployed;
- where they were running;
- which systems depended on them;
- which data moved between them;
- who owned each component;
- how it was monitored;
- and what happened when it failed.
Each conversation added another piece.
One engineer described an application dependency. Another explained the infrastructure beneath it. Someone else identified an operational process that existed only because of the way two systems interacted.
Gradually, several architectural views began to emerge:
- Infrastructure view: hosts, networks, storage, regions and deployment locations.
- Platform view: shared services, databases, messaging, deployment and operational tooling.
- Application view: services, responsibilities and functional relationships.
- Data-flow view: how information moved through the environment.
- Failure-domain view: which components could fail together and which services would be affected.
These were not separate architectures. They were different ways of examining the same system.
The scale surprised everyone
When the complete picture finally came together, even I was surprised.
The platform consisted of approximately 1,000 machines distributed across three continents.
Global Platform
|
+--------------------+--------------------+
| | |
v v v
Continent A Continent B Continent C
| | |
Applications, Applications, Applications,
infrastructure, infrastructure, infrastructure,
databases and databases and databases and
platform services platform services platform services
Approximately 1,000 machinesMany people were genuinely shocked when I shared the consolidated view.
They had been operating parts of the platform every day without recognising the full scale and complexity of the environment they collectively supported.
This was not because they lacked competence.
The organisation had structured knowledge around teams and responsibilities. No one had been given clear ownership of the complete architectural picture.
Architecture is not one diagram
A single diagram cannot adequately describe a complex platform.
An executive may need a high-level view showing major capabilities and regions. An engineer investigating latency needs dependencies and data flows. A security specialist needs trust boundaries. An incident manager needs failure domains and operational ownership.
The architectural picture should therefore consist of connected views at different levels of detail.
A useful hierarchy might look like this:
Level 1: Business and platform context
|
Level 2: Major systems and regions
|
Level 3: Services, data flows and dependencies
|
Level 4: Deployment and infrastructure
|
Level 5: Component-specific technical detailEach level should connect to the next.
The high-level view provides orientation. The lower-level views provide evidence and operational detail.
Visibility changed the conversation
Once the architecture became visible, improvement opportunities that had previously seemed isolated began to connect.
It became possible to discuss:
- unnecessary complexity;
- duplicated services;
- weak or missing ownership;
- observability gaps;
- fragile dependencies;
- large failure domains;
- manual operational work;
- and avoidable downtime.
Before the shared model existed, each problem could be treated as a local issue.
Afterward, the organisation could see that several problems were consequences of the same architectural relationships.
Reducing unnecessary complexity
Complexity is not automatically bad.
A global platform may genuinely require:
- regional deployments;
- redundancy;
- several data stores;
- asynchronous processing;
- and specialised infrastructure.
The problem is complexity that no longer has a clear reason.
Once the architecture was documented, components could be challenged constructively:
- Does this service still have a consumer?
- Why does this region have a different implementation?
- Could these two components be consolidated?
- Is this dependency intentional?
- Does this operational process compensate for an avoidable design problem?
Without the shared picture, removing a component feels risky because its relationships are unknown.
Visibility creates the confidence required to simplify.
Improving observability
Observability should follow the architecture.
If the organisation does not understand the service boundaries and dependencies, it may collect large quantities of telemetry without measuring the behaviour that actually matters.
The architectural model helps connect:
User outcome
|
Service-level indicator
|
Application behaviour
|
Platform dependency
|
Infrastructure resourceAn alert can then be associated with an owner and an expected response.
A dashboard can answer a defined question.
Telemetry becomes part of an operational model rather than an accumulation of available measurements.
Understanding failure domains
A component-level view may show that each service has redundancy.
The wider architecture may reveal that several supposedly independent services share:
- the same database;
- the same network path;
- the same storage;
- the same deployment mechanism;
- or the same operational dependency.
That shared dependency represents a larger failure domain than the individual service diagrams suggest.
Making it visible allows the organisation to decide whether the risk is acceptable or needs to be reduced.
The human resistance to clarity
What disappointed me was discovering that not everyone welcomed this increased visibility.
Some people appeared more comfortable with the earlier environment, where knowledge remained fragmented and particular individuals were required whenever something failed.
I interpreted part of this resistance as attachment to hero behaviour: being the person who saves the day because only that person understands a particular corner of the system.
That interpretation may not explain every reaction. People can resist architectural documentation for many reasons:
- fear of judgement;
- lack of time;
- concern that documentation will become outdated;
- loss of autonomy;
- or previous experiences with documentation exercises that produced no practical benefit.
Nevertheless, knowledge concentration can provide status and organisational security.
When only one person understands a system, that person appears indispensable.
The organisation, however, becomes fragile.
Heroics do not scale
Individual expertise is valuable.
The problem begins when the operating model depends on emergency intervention from specific people.
A hero-based environment often has several characteristics:
- knowledge is undocumented;
- routine work requires escalation;
- incidents repeatedly reach the same people;
- manual recovery is celebrated;
- root causes remain unresolved;
- and operational success depends on personal availability.
This creates immediate rewards for rescue behaviour while providing little incentive to remove the conditions requiring rescue.
A scalable system works differently.
Knowledge is shared. Responsibilities are explicit. Recovery procedures exist. Repeated work is automated. Incidents lead to structural improvement.
The objective is not to eliminate expert engineers.
It is to use their expertise to make the organisation less dependent on emergencies.
Architecture needs ownership
An architectural picture will not maintain itself.
Someone must be accountable for ensuring that the views remain useful.
Ownership does not mean that one architect personally documents every component. The engineers closest to each area should maintain the relevant detail.
The architectural owner provides:
- standards;
- structure;
- integration between views;
- review;
- and accountability for the complete model.
A federated approach can work well:
Central architecture ownership
|
+----------+----------+
| | |
Team A Team B Team C
maintains maintains maintains
its view its view its view
\ | /
+---- Shared model --+The organisation gains a coherent system view without separating documentation from the teams implementing and operating the platform.
Documentation should be part of change
Architecture becomes outdated when documentation is treated as a separate project.
A more sustainable approach connects documentation to ordinary engineering work.
Changes affecting components, dependencies, data flows or deployment should include updates to the relevant architectural view.
Reviews can ask:
- Does this introduce a new dependency?
- Does it change a trust boundary?
- Does it increase a failure domain?
- Does ownership change?
- Has the architecture been updated?
The documentation then evolves with the platform instead of attempting to reconstruct it years later.
Start with what is useful
An organisation with little architectural documentation should not begin by attempting to model everything perfectly.
A practical sequence is:
- Identify the major systems and regions.
- Document the most important user and data flows.
- Map critical dependencies.
- Assign ownership.
- Identify important trust and failure boundaries.
- Add detail where it supports an operational or design decision.
The purpose is understanding, not diagram completeness.
A simple, accurate model is more valuable than a visually elaborate model nobody trusts.
Once the system is visible
Creating the architectural picture did more than improve documentation.
It changed conversations.
Teams could discuss the same system instead of defending separate interpretations. Priorities became easier to explain. Operational risks could be connected to architectural causes. Simplification opportunities gained visible evidence.
The platform had existed before the diagrams.
What changed was the organisation’s ability to reason about it collectively.
That is the value of architecture.
It is not merely a diagram produced for governance or presentation. It is a shared model that helps people understand what they have built, how it behaves and where improvement is possible.
When nobody owns the complete picture, complexity grows in the gaps between teams.
Once the system becomes visible, improvement becomes possible.
This article is based on my original ideas, experience, analysis and conclusions. Artificial intelligence tools were subsequently used as editorial and research assistants to review grammar and wording, improve structure and presentation, organise some arguments into clearer logical sections, and help review references to legal, regulatory and technical concepts.
Where relevant, factual and regulatory references were checked against the sources cited in the article. AI assistance does not replace professional legal, regulatory, financial or technical advice, and the final selection, interpretation, opinions and conclusions presented here remain my own.
Comments
Post a Comment