Article
One body, many masters: DevOps is a culture problem with an architectural cause
Somewhere between the coining of the term and the present day, an idea invented to dissolve a boundary was turned into an organisation chart.
DevOps is a compound word. Development and operations. It exists because the two had been separated, and because the separation was doing visible damage: developers built software and threw it over a wall, operations caught it and were made accountable for a thing they had no hand in shaping. The word named the remedy — that the people who build a thing should be the people who run it.
Today it is a job title. It is a department. There are DevOps engineers, DevOps teams, heads of DevOps, and in larger organisations a DevOps function with its own cost centre, its own director and its own box on the slide. Recruiters have a dedicated pipeline for it. Which means the industry took a word coined to remove a wall and used it to build one, then hung the name of the remedy on the door.
I made a version of this argument three years ago and concluded that it was a symptom of an ailing engineering culture. I still hold that. What I would add now is the other half of it, because treating culture alone has a poor record and the reason is not mysterious.
Culture and architecture are two views of one fault. What people optimise for, what they refuse to own and who they blame at three in the morning are shaped by lines that were drawn long before any of them were hired — and those lines are enterprise architecture, whether or not anybody with that title drew them. But the traffic runs both ways. Once a structure has been in place for a few years it produces beliefs, and the beliefs then defend the structure: that developers cannot be trusted with production, that operations exists to say no, that this is simply how a company of our size and our regulator must work. By then, redrawing the boundary is not enough on its own, because the people on either side of it will faithfully reconstruct the old wall inside the new shape.
So it has to be worked from both ends. Values without structural change are a poster. Structural change without the values work is a reorganisation that reverts within the year. What follows is mostly about the end that gets neglected, which is the architecture — not because it matters more, but because almost every attempt at this begins and ends with the culture.
A body with several masters
Consider what a human being would be if each part answered to a different authority.
One person has the hands and is measured on grip strength. Another has the feet and is measured on distance covered. A third holds the eyes and is rewarded for detail observed. A fourth has the arms and is judged on reach. Nobody is negligent; nobody is malicious. Each is diligent, and each optimises the thing they were given and can be judged on.
The body does not go anywhere. The feet advance while the eyes are examining something behind. The hands grip an object the feet are already walking away from. Every limb hits its number. The organism starves.
What makes a body work is not that each part is well managed. It is that there is one nervous system carrying a single intent, and one circulatory system carrying a shared supply. Specialisation is not the problem — a body is nothing but specialised organs. The problem is specialisation without a common command, and the moment you cut the nerve, competence in the limb stops helping and starts hurting, because a strong limb pulling the wrong way does more damage than a weak one.
That is what a business does to itself when it splits building from running.
What the split actually does
The damage is not vague. It is arithmetic, and it is written down in two documents that are rarely read side by side.
The delivery organisation is measured on throughput: features shipped, roadmap delivered, dates met. The operations organisation is measured on stability: uptime, incident count, change failure rate. Both are legitimate. Both are, in a system where a change is the only way to deliver value and also the only meaningful source of risk, in direct opposition. One group is paid to increase the rate of change and the other is paid to reduce it.
Nobody in that arrangement is behaving badly. The change advisory board that meets on Thursdays and returns your release with a question is doing exactly what it was constituted to do. The team that ships around it, or wraps a schema change in a feature flag so it does not count as a change, is also doing exactly what it was measured to do. The friction everybody complains about is not a failure of goodwill. It is the intended output of the design.
Then there is the incident. Ten people from ten departments on a bridge call, each arriving with a partial view and a private need to establish that the fault lies elsewhere, spending the first forty minutes on jurisdiction before anyone touches the system. This is universally described as a cultural failure, and it is one — the defensiveness is real and it is corrosive. But it is also, and more usefully, a structural one: no single party holds enough of the picture to diagnose the fault, and every party holds enough of the blame to be defensive. Put the same ten people in one team with one objective and the call takes four minutes, and the culture on that call improves without anyone having addressed it directly.
There is a cost nobody puts in a budget line, which is the knowledge that never forms. Feedback from production is the only real information a system produces about itself. When it lands on people who cannot change the code, and the people who can change the code never hear it, the organisation stops learning about its own product. It does not degrade suddenly. It just quietly ceases to improve, and the reason is invisible on every dashboard anyone is looking at.
Werner Vogels put the alternative more concisely than I can, and it has been quoted for a decade without being acted on:
Giving developers operational responsibilities has greatly enhanced the quality of the services, both from a customer and a technology point of view. The traditional model is that you take your software to the wall that separates development and operations and throw it over and then forget about it. Not at Amazon. You build it, you run it. This brings developers into contact with the day-to-day operation of their software. It also brings them into day-to-day contact with the customer. This customer feedback loop is essential for improving the quality of the service.
Note what he is describing. It is not a tooling decision and not a values statement. It is an ownership boundary — an architectural choice about where responsibility for a system begins and ends.
Why the architecture is the half that gets missed
Conway observed sixty years ago that an organisation ships its own communication structure. The observation is usually quoted as a warning about software. It is better read in the other direction: if you know what system you need, you already know a great deal about the structure that can produce it, and every structure you choose instead is a decision to ship something else.
That makes the org chart a technical artefact, and drawing it a technical decision — which is precisely why it should not be made purely on the basis of headcount, procurement convenience or which vendor happened to sell the tooling. When responsibility for a service is divided across four departments, the interfaces between those departments become interfaces in the system. They acquire queues, ticket forms, service level agreements and waiting time. That handoff cost is real engineering cost, and it is paid every day, forever, by everyone — but because it is distributed across four budgets, no one department ever sees the total, and so no one department can make the case for removing it.
This is the specific work of enterprise architecture, and it is why the discipline is not a diagramming exercise. Enterprise architecture is the practice of keeping business intent, the systems that deliver it, and the people accountable for them aligned to one another — deciding where boundaries fall, what each side of a boundary owes the other, and who holds the decision.
Get it wrong and no amount of values training or team-building survives contact with the payslip, because you are asking people to behave against their own measurement and they will not do it for long. Get it right and you have not fixed the culture either — you have made the good culture affordable. Cooperation stops costing an individual something to offer, which is the condition under which the values you have been talking about for two years finally start to hold. The structure does not create the culture. It decides whether the culture you want is sustainable or heroic, and heroism does not scale past the people currently performing it.
Hiring a DevOps department is what an organisation does when it recognises the symptom and treats it at the wrong layer. The wall is causing pain, so a team is created to sit on top of the wall and pass things across it more efficiently. The wall remains. It now has a dedicated staff, a budget, a director with a career interest in its continuation, and a name that makes it very difficult to discuss.
Specialisation is not the enemy
None of this argues that every engineer should do everything, and the reading that it does is the reason many attempts at this fail badly. A body is nothing but specialists. You want people who know Kubernetes deeply, people who know networks, people who know databases at a level a product team never will. Cutting those people out and asking a product engineer to absorb their work is not integration — it is amputation with extra steps.
The distinction that matters is not whether a specialist team exists. It is what that team owns.
A platform team that builds and runs a product — a paved road that other teams consume, with an interface, a version, and users who could in principle refuse to use it — is an organ. It is inside the nervous system. Its success is measured by whether the teams using it go faster.
A team that owns a stage in someone else’s sequence — that receives work, processes it, and passes it on, whose approval is required rather than chosen — is a wall, whatever it is called. Its success is measured by the throughput of its own queue, which is a number that can improve while the business gets worse.
Same skills. Same people, often. Entirely different architecture, and entirely different outcomes.
Five questions worth asking
If you want to know which one you have, the org chart will not tell you. These will.
- Who is paged when the service you shipped last month breaks at three in the morning? If it is not someone who can read and change the code, you have a wall.
- Can a team put a change into production without a person outside that team saying yes? Automated policy is not a person. A gate held by another department’s queue is.
- Are two of your scorecards in direct opposition? Find the one measured on rate of change and the one measured on stability. If they are held by different leaders with different budgets, the conflict is designed in, and it will express itself as personality.
- When you count the cost of a handoff — the ticket, the wait, the re-explaining, the rework — whose budget does it land in? If the answer is nobody’s, it will never be removed.
- Does your definition of done include production? If a story can be complete before a customer can use it, then every measurement downstream of that definition is measuring something other than value.
Bringing a body back under one command
The remedy is not a reorganisation, and it is certainly not a renaming. It is architecture, done in a specific order: establish what the business is actually trying to do, identify the capabilities that deliver it, draw the service boundaries so that each one can be owned end to end by a team that can be held accountable for it, and then — only then — decide what platform, tooling and specialist support that structure needs in order to work. Ownership first, technology second. Almost every failed transformation does it the other way round.
And then the other half, which cannot be drawn on a diagram and cannot be skipped. Ownership has to be accepted as well as assigned, and a team handed responsibility for production on Monday does not feel accountable for it by Friday. That is the work of a definition of done that ends in production and is enforced; of engineers carrying the pager for their own service and leaders visibly carrying it with them; of incident reviews that establish cause rather than fault, held often enough to be ordinary; of hiring and promoting for the willingness to own a thing end to end rather than for depth in a silo; and of stopping the quiet reward of heroics, because a business that celebrates the rescue is buying more of what needed rescuing. None of that works while the structure contradicts it — and the structure, once corrected, does not fill itself in.
That is the work we do at Mamluk. Our enterprise architecture practice exists for exactly this class of problem: target-state and transition architectures that account for the operating model as well as the technology, service and ownership boundaries that hold under load, architectural governance that puts decisions where the knowledge is, and honest second opinions on programmes already in flight. We are engineers who run our own cloud and our own Kubernetes distribution, so the boundaries we draw are ones we have had to live inside at three in the morning.
If you are hiring for DevOps, the vacancy is real and the pain behind it is real. But before you fill it, it is worth asking what the role is compensating for — because a body with several masters does not need a better coordinator between the limbs. It needs a nervous system.
— Meezaan-ud-Din Abdu Dhil-Jalali Wal-Ikram, founder of Mamluk