Conway's Law

(Originally drafted in July 2019 during my tenure at Xero. Republished here in October 2026 with updated context and removed proprietary references)
“A system is never the sum of its parts. It is the product of the interactions of its parts.” — Dr. Russell Ackoff
Every sufficiently large software platform gives its users the impression that it’s a single, unified thing running on unicorns and magic. In reality, it has always boiled down to software running on servers that is developed and managed by people. And those people - how they communicate, where they sit, who owns what - quietly shape the systems they build in ways they rarely notice.
That’s not an accident. That’s Conway’s Law.
Conway’s wha?
“Organizations which design systems… are constrained to produce designs which are copies of the communication structures of these organizations.” — Melvin Conway
Conway’s observation, from way back in 1967, is that in order for a software module to function, multiple authors must communicate frequently with each other. Therefore, the interface structure of a system reflects the social boundaries of the organisation that produced it – and communication across those boundaries is harder than communication within them.
Put bluntly: if your team structure looks like a mess, your architecture will too.
It’s connections all the way down
One way to read the history of technology is as a history of connecting people together. Google connects billions of people to data. Facebook and LinkedIn connect people to each other. Uber connects riders to drivers. Amazon and eBay connect buyers to sellers. Slack connects employees into shared workspaces. GitHub connects developers to codebases. Morse code connected hundreds of thousands of people through a telegraph network and was arguably the precursor to the civilization-changing global Internet.

All of these effectively model graph networks, and the most successful platforms of the last two decades are the ones that built their business around facilitating connections. Even at the infrastructure level, this insight holds: Amazon’s first AWS service, launched in 2006, was the Simple Queue Service – born out of the need to decouple systems so they could operate independently of each other but still share data.
How did an organisation as enormous as Amazon achieve that kind of delivery velocity? Surely a large, interconnected organisation should be bottlenecked by its own processes and move slower than smaller ones?
Maybe we need to look at how they were structured – both at the organisational level and at the system level – to find out.
How monoliths actually happen
Nobody sets out to build a giant tangled monolith. It happens incrementally, each step perfectly rational at the time.
In my experience, it goes something like this. A product starts life as a single codebase serving one purpose, often modelling one core relationship between users. That’s the goal: achieved. Then someone realises users don’t just have a one-to-one relationship with the product but a many-to-many relationship with each other – so a new codebase gets split out, sometimes literally by copying and pasting. Then billing needs its own system. Then an API to let third parties leverage the data. Then an acquisition brings in a payroll product, a project management tool, a document-scanning service, each in its own codebase, each developed in its own location, by teams who never needed to share information.
Individually, each of these systems serves its purpose. Collectively? They become a distributed system that nobody designed. And they all want data from the same place, so they ‘reach in to’ and ‘pull out’ whatever they need – typically from one central database that has become the go-to for everything.
Want to get an invoice? Hit the main database. Want to get reports? Hit the main database. Want to get any data at all? Hit the main database.
To enable systems to notify one another that something has changed, polling mechanisms get bolted on: one system constantly asks another whether there are any changes, and that system sticks a micro-API endpoint in front of the database and polls it until the end of time.
As you can imagine, for performance, this is bad.
And here’s where Conway’s Law bites: each system was historically built by an independent team, organised around its own domain, in its own office, in its own time zone. The system boundaries are the communication boundaries. When I mapped out the architecture of the platform I worked on, the correlation between where code lived and where teams lived was impossible to miss.
The coupling trap
Shared code promises time-to-market wins: why rewrite the accounting engine when importing a NuGet package already has all the functionality you need? That makes perfect sense on a small scale, but scaled up across a large organisation the problems snowball: shared ownership, undocumented breaking changes, security holes, overlapping bug fixes.
While working on one product I discovered we were pulling in over 250 internal NuGet packages, slowing builds and delivery times. Investigating the individual method calls revealed that the majority of the functionality fell into four distinct buckets: logging, monitoring, login integration, or accounting logic.
And that last bucket was telling – because nearly every product in the company needed the same accounting data. One product had its own journal line reader. So did the tax product, and the dashboard, and the reporting engine, and payroll. Everyone needed journal lines because journals are integral to the core domain.
When coupling becomes so densely tangled that you have to surgically unpick dependencies just to upgrade a third-party library in a downstream system - all while the airplane is in flight and you’re still serving existing customers – you have an architectural problem.
Horizontal vs. vertical slices
Most legacy architectures are horizontal: capabilities are shared between product groups, organised around technical layers. The moment you want to do anything outside your slice, it hurts.
The alternative is a vertical slice architecture, where the organisation is structured around business domains and bounded contexts, following Domain-Driven Design:
- Vertical slice architecture looks at your domains and bounded contexts
- Once the domains are known, the organisation is structured around those domains
- Interactions between domains become clearly defined
- Domain events raised by one domain can be consumed by another
This is exactly what the rise of Micro Frontends in the front-end world reflects: teams formed around vertical slices of business functionality rather than technical capabilities, with each slice owned end-to-end by a single team.
Even then, horizontal and vertical might just be visualisations that help humans reason about architecture. The reality is always messier: systems call systems that call other systems - a massively distributed system whether you designed it that way or not. The explosion of Service Mesh software confirms this, attempting to combat the classic Fallacies of Distributed Computing. Maybe that complexity is intrinsic to how modern cloud software has to work.
Event-driven architecture: decoupling by talking less
“In an event-driven architecture, components perform activity in response to receiving events and emit events to trigger activities in other components.” — Nat Pryce
Whatever architecture style you choose, your components still need a way to communicate so the real-world interactions in your domain can continue to be modelled. The mistake most platforms make is relying exclusively on pull-based mechanisms - polling - instead of investing in push-based ones.
In an event-driven architecture, systems no longer poll each other for changes; they’re told when changes happen. The pressure on the overall architecture drops, systems become genuinely decoupled, microservices can be replaced individually, and the system can adapt to the domain it serves.
This is where the famous Bezos API mandate points. Reportedly issued around 2002, it demanded that all teams expose their data and functionality through service interfaces, communicate only through those interfaces – no direct linking, no direct reads of another team’s data store, no back-doors - and that every interface be designed from the ground up to be externalisable. It doesn’t take much hindsight to see how that decision worked out.
Notice what that mandate really is: Conway’s Law used deliberately. Instead of letting the organisation’s communication structure accidentally constrain the design, you reshape the communication structure to produce the design you want.
The payoff: automation
Why go to all this trouble? Because decoupled, event-based architecture is what makes real automation possible.
Imagine taking a photo of a document on a phone in the middle of a field somewhere, and the upload triggers automated processes that scan the photo, match it to a transaction, reconcile that transaction line with a bank feed, and push a notification back to the phone when it’s done – no human intervention required.
Imagine an onboarding process where the first step asks you to upload one of your existing documents, and the system extracts your company name, logo, address, bank details and industry from it – leaving you merely to verify the results rather than type them in. What a first impression that would be.
And here’s the part people miss: those little CAPTCHA components that ask you to pick out crosswalks and fire hydrants? You might think you’re verifying you’re human, but you’re simultaneously making a machine learning model more accurate. Every user interaction in a well-designed system is training data.
For all of that to work, though, the underlying systems need to talk to each other through standardised events – not polling, not one-off bespoke integrations.
Closing thoughts
When I started writing this post I had quite a few ideas floating around and this all kind of ended up as a mish-mash of words that I’ve slowly massaged into something coherent. So I figured the best way to close would be a few succinct bullet points:
- Large-scale architecture is almost always historical: it evolves to where it is today through startup-era pragmatism, acquisitions, and physical expansion, not deliberate design.
- Conway’s Law appears to have a direct correlation with internal architecture when systems are built as independent, siloed codebases by teams that never had to share information.
- Over time, those systems become increasingly coupled, and stability and uptime suffer as a result.
- To move faster and deliver more value, we need decoupled systems that can be released independently with minimal knock-on effects.
- Systems can only be truly decoupled when they have an agreed, standardised messaging mechanism and format. A smattering of SQS/SNS here, Kinesis there, webhooks somewhere else isn’t a strategy.
- Distributed systems introduce a whole host of new complications – latency, availability, security, observability – that you need to be aware of, but they’ve been proven to work at scale for financial-grade systems.
One more thing worth pondering. Organisations typically restructure around their customers. If Conway’s Law holds, and your architecture mirrors your organisational structure, then your architecture will follow your customers’ communication structures too. That’s either a very convenient accident or a design principle hiding in plain sight.
Related Posts
2026-09-27
2024-08-30
2024-03-03