Based on a LinkedIn post originally published on 8 January 2026
During the months leading into 2026, I began deliberately shifting part of my professional focus back towards hands-on technical work.
One element of that process was preparing for the AWS Certified DevOps Engineer – Professional certification, using Pluralsight as my principal learning platform.
The certification provided structure, but obtaining another credential was not the real objective.
My goal was to revisit the engineering depth behind modern DevOps practices: continuous integration and delivery, infrastructure as code, observability, resilience and secure automation at scale.
These were not unfamiliar subjects. I had worked with them for years through engineering leadership, architecture, platform operations, security and organisational transformation. However, understanding an area strategically is different from implementing it directly.
I wanted to strengthen that implementation muscle again.
Certification is evidence of study, not proof of mastery
Professional certifications can be useful. They establish a syllabus, organise a large subject into manageable areas and provide a measurable objective.
They can also become misleading when passing the examination becomes the primary goal.
Memorising service names, interface options and recommended answers is not equivalent to understanding how a system will behave in production. A certification cannot prove that someone can diagnose an ambiguous failure, design a safe deployment process or decide which architectural trade-off is appropriate for a particular organisation.
The examination may ask which AWS service satisfies a stated requirement. Real engineering begins when the requirements are incomplete, contradictory or still changing.
That is why I treated certification preparation as a framework for deeper practical review rather than as an end in itself.
For each subject, the more valuable questions were:
Why does this practice exist?
Which failure is it intended to prevent?
Under what conditions does it stop working?
What operational burden does it introduce?
How would I implement and test it?
What evidence would demonstrate that it works?
How would another engineer understand and maintain it?
These questions move the exercise beyond examination preparation and towards engineering competence.
Revisiting continuous delivery
CI/CD is frequently reduced to the presence of a pipeline.
A repository contains a configuration file, an automated job compiles the software and another job deploys it. The organisation can then claim to practise continuous integration or continuous delivery.
The existence of automation, however, says little about its quality.
A useful delivery process must address questions such as:
Which changes trigger the pipeline?
What is tested at each stage?
Are builds reproducible?
How are artefacts identified and promoted?
Can the same artefact move safely between environments?
How are credentials protected?
What approvals are genuinely required?
How are failed deployments detected?
Can a release be stopped, rolled back or rolled forward safely?
What evidence remains after deployment?
A pipeline should not merely automate commands. It should encode a controlled and understandable delivery process.
This is where hands-on work matters. Architecture diagrams and strategy documents can describe an intended flow, but implementation reveals whether the design is practical.
Infrastructure as code is more than automation
Infrastructure as code is another area in which the name can create a false sense of maturity.
Representing infrastructure in a declarative file is only the beginning.
The real benefits emerge when infrastructure definitions are:
version controlled;
reviewed;
tested;
repeatable;
modular where appropriate;
protected against uncontrolled changes;
and connected to a disciplined delivery process.
A poor infrastructure-as-code repository can reproduce mistakes more efficiently without making the environment easier to understand or operate.
The important questions include:
Who owns each component?
How is state protected?
How are environment differences handled?
How are modules versioned?
How are destructive changes identified?
What prevents manual configuration from creating drift?
How are emergency changes reconciled afterward?
Can the environment be reconstructed from the declared configuration?
Again, these questions are difficult to explore through theory alone. Building and changing real environments exposes the consequences of the design.
Observability must support decisions
Modern cloud platforms can generate enormous quantities of metrics, logs, traces and events.
Collecting that information is comparatively easy. Turning it into operational understanding is much harder.
Observability should help engineers answer concrete questions:
Is the service functioning correctly?
Are users experiencing a problem?
Which dependency is failing?
When did the behaviour change?
What was deployed before the change?
Is the system approaching a capacity limit?
Which alert requires immediate action?
Which signal is merely noise?
An organisation can have sophisticated dashboards and still lack operational visibility.
This happens when instrumentation grows without clear ownership or purpose. Metrics are collected because they are available. Alerts are created because thresholds can be defined. Dashboards multiply, but nobody knows which one represents the authoritative view of service health.
Hands-on implementation forces the engineer to connect telemetry to system behaviour and operational decisions.
Resilience must be tested
Cloud-native architecture often includes language such as “highly available,” “fault tolerant” and “self-healing.”
These descriptions should never be accepted solely because a managed service or architectural pattern is present.
Resilience is a property of the complete system.
A service may run across multiple availability zones while still depending on a single external system. A database may be replicated while the application cannot reconnect after failover. A deployment may use multiple instances while a shared configuration error disables all of them simultaneously.
The relevant questions are practical:
What failures are expected?
Which failures are automatically handled?
How long does recovery take?
Is data lost during recovery?
Which components still require manual intervention?
What happens when several failures occur together?
Has the recovery process actually been exercised?
A resilient design must be verified by deliberately testing failure scenarios. Assumed resilience is not resilience.
Secure automation at scale
Automation amplifies both good and bad decisions.
A secure automated process can apply consistent controls across thousands of resources. An insecure one can distribute excessive permissions, exposed credentials or incorrect configuration just as efficiently.
Secure automation requires attention to:
identity and access management;
separation of duties;
credential storage and rotation;
least-privilege permissions;
audit trails;
policy enforcement;
dependency integrity;
artefact provenance;
and controlled emergency access.
Security should be part of the delivery mechanism rather than an external review added after implementation.
The objective is not to place a manual security gate in front of every change. It is to encode appropriate controls into the engineering process so that secure behaviour becomes the normal and efficient path.
Building, breaking and fixing
What I appreciated most about returning to practical study was its honesty.
A system either behaves as expected or it does not.
A deployment succeeds, fails or produces an ambiguous state. An alert identifies a real problem or creates noise. A recovery procedure restores the service or exposes an assumption that was never tested.
There is no shortcut around building, breaking and fixing things.
This feedback loop is essential because it converts abstract understanding into judgement. Each implementation decision creates evidence. Each failure challenges an assumption. Each correction improves both the system and the engineer’s mental model of it.
Combining leadership with implementation
Returning to hands-on work did not mean discarding leadership or architectural experience.
The objective was to combine them.
Leadership provides an understanding of organisational constraints, priorities, team dynamics and long-term consequences. Architecture provides a view across systems and dependencies. Hands-on engineering keeps those decisions grounded in implementation reality.
The strongest technical leadership should remain connected to the work closely enough to understand its real cost and complexity.
Similarly, experienced hands-on engineers can contribute more effectively when they understand why priorities exist, how risks are evaluated and how local decisions affect the wider organisation.
These capabilities reinforce one another.
Staying technically honest
Technology changes continuously. Cloud services evolve, tools are replaced and recommended practices are revised.
Some principles remain remarkably stable:
automate repeatable work;
keep changes small and observable;
design for failure;
protect identities and data;
make ownership explicit;
reduce unnecessary complexity;
and verify assumptions with evidence.
A structured certification programme can help identify what has changed. Practical implementation reveals what still holds true.
That was the real purpose of this journey: not simply to add another line to a CV, but to reconnect strategy with execution and remain honest about the depth of my own understanding.
In engineering, familiarity is not the same as fluency.
Fluency comes from doing the work.
This article is based on my original ideas, experience, analysis and conclusions. Artificial intelligence tools were subsequently used as editorial and research assistants to review grammar and wording, improve structure and presentation, organise some arguments into clearer logical sections, and help review references to legal, regulatory and technical concepts.
Where relevant, factual and regulatory references were checked against the sources cited in the article. AI assistance does not replace professional legal, regulatory, financial or technical advice, and the final selection, interpretation, opinions and conclusions presented here remain my own.
Comments
Post a Comment