Principles of Distributed Systems

How to build a system that does not suck

Preface

Why are we designing distributed systems?
Why are we doing so-called “microservices”?

There’s only one right answer: Independent deployability.

We want to deploy our applications independently, but from what — or whom?
“From other applications, of course!” — you might think. Although this seems like a straightforward answer, it’s not the whole truth.
The more appropriate answer is: “From other people.”

Let me explain. Or rather give you an example.

Some time ago I worked on a project where we had ~17 “agile” teams working on a single codebase. It was a classic monolithic Java application rooted in the early 2010s. When I joined the project, it was already a mess.
First, the application was deployed only 4 times a year(!) ± urgent bug fixes — unheard of in the modern world.
Second, the deployment process was complicated and was taking ~2 days(!).
The application (huge e-commerce website) was going down every quarter on Saturday evening and was up and running on Monday morning. Over a hundred people were involved in the deployment process — QAs, Engineers, and Managers from every possible affected system. There was a rumor that fax machines were involved. I still haven’t figured out the main reason for such rare deployments — was it a cost factor, or the bureaucracy?

And that’s not all.

The development process was also far from perfect.
We had 4-week sprints. We’re agile of course ;)
At the end of every sprint, the new release was forked from master branch. This means that if a team wants its work to be a part of the release, they should merge their changes into master before the release is forked. And since every team wanted their work to be in the release, every single team was merging their changes into master branch at the end of every sprint. Multiplying this by 17 teams (~100 devs), a single codebase, and 4 weeks of development — you can imagine the mess.
To deal with the mess, a code freeze was introduced. At the end of every sprint, the last two days were declared a code freeze. The release branch was forked from master and the rest of the sprint was dedicated to resolving merge conflicts, testing, stabilizing the code, and fixing newly created bugs.
Very interesting process, and very toxic.
Those days made “moving a .war archive via FTP” look like modern CI/CD.

The codebase was all over the place, too, of course.
That is why we had 4-week sprints — it was physically impossible to build anything tangible in 2 weeks, and so management doubled the sprint length. Nowadays, this is a very common red flag… or should I say pronouncement of death? But I digress.

This is exactly why “microservices” were born — out of frustration.
“This can’t continue. Let’s cut the part of the application we are responsible for to avoid dealing with others.” — isn’t this a lovely statement?
Of course, the first attempts went terribly wrong.
Even more complexity was introduced:

  • How to build, deploy, and monitor this zoo?
  • Instead of invoking a method, we should make an HTTP call, right?
  • How to deal with transactions?
  • I need to change a service we don’t own now, but the owner has no capacity for the next few months. WTF?
  • Why has the cost exploded?
  • Yadda yadda yadda…

Eventually, the idea evolved, matured, and got shaped into a simple principle of Independent deployability.

Unfortunately, we can’t have a cake and eat it too. Egocentric engineers jumped on the “microservices” bandwagon like it was a saving grace. They wanted to escape horrible monoliths and live happily in the beautiful world of microservices, whatever that meant. I should note, that those are the same developers who made the monoliths horrible in the first place… Humph.

The following principles should provide a starting point to build an independently deployable system, prevent excessive over-engineering, and allow for innovation.

Principles

Customer First

It all starts with the customer. Always. Whether it’s the end-user paying for a service, the back-office team managing operations, or a third-party system pulling data via API — they’re all customers. And every decision we make, every line of code we write, needs to serve them. But here’s the tricky part: defining “customer” is often half the battle. You can’t build something great until you truly understand who you’re building it for.

With the Customer First principle, the customer’s needs define everything — even your domain boundaries. These boundaries shouldn’t be driven by how the system is implemented but by how the customers interact with it. And let’s be clear: one-size-fits-all doesn’t work. Separate the concerns. Address each customer’s need individually, cleanly, and effectively. Yes, there will be trade-offs. Yes, there will be technical constraints. But if you’re not serving the customer — you’re doing it wrong. Keep it simple. Keep it focused. And always ask yourself: Does this serve the customer?

Share Nothing

It’s only after you’ve lost everything that you’re free to do anything. Nothing is static, everything is evolving, everything is falling apart. — Fight Club

The Share Nothing principle is simple: independence equals freedom. A system that depends on nothing can’t be held back by anything. It’s the cornerstone of resilient distributed systems. The less you share, the fewer dependencies you have. And the fewer dependencies you have, the easier it is to deploy, evolve, and fix your system without breaking someone else’s — or worse, waiting for someone else to fix theirs.

Now, don’t kid yourself. Total independence is a pipe dream. Every system has dependencies. But the fewer there are, the better. Don’t share a database! That’s a bad rap, and everyone knows it. One schema tweak and bam! Everyone downstream is in a world of hurt. But what about REST APIs? Same deal. It’s just another form of coupling. Someone changes the contract and your system breaks. So, what’s the answer? Share as little as humanly possible. Make your system stand on its own two feet.

If you must share, do it carefully. Define boundaries. Version APIs. Use patterns like CQRS (Command Query Responsibility Segregation), events, or queues to decouple workflows. Independence doesn’t mean isolation — it means freedom. And freedom is the whole point.

Event-Driven First

In the presence of a network partition, the system can be either strongly consistent or highly available, but not both. — CAP theorem as phrased in Designing Data-Intensive Applications by Martin Kleppmann

The Event-Driven First principle is the engine of the design: when you’re building distributed systems, events are king. Why? Because the CAP Theorem doesn’t care about your deadlines or uptime guarantees. In a network partition, you’ve got two choices: stay consistent and freeze, or stay available and keep moving. it says: keep moving.

Here’s the reality: your system has to communicate. It has to share data. But here’s what it shouldn’t share — faults and failures. When one piece goes down, the rest of the system should stay up. No dragging everyone into the mud. No expanding the blast radius. We want resilience. We want loose coupling. We want independence.

How do you get there? Asynchronous communication. Stop tying services together with synchronous, tightly coupled calls. Use events to decouple producers and consumers. When one service crashes or slows down, the others keep humming along. Yes, you’ll sacrifice strong consistency. But let’s be real — perfect consistency is a myth in distributed systems. What you gain is high availability, scalability, and the ability to recover gracefully when everything goes sideways.

The bottom line? Build your system around events. Let services talk to each other without depending on each other. Stay alive, stay fast, and stay independent. That’s the principle in action.

Minimize Costs

Minimize Costs starts with simplicity, not spreadsheets: every thing you introduce — a service, a queue, a database, a framework — is one more thing to navigate. To understand, to patch, to update, to debug at 2 AM. And let’s face it, maintaining things isn’t fun. Keeping pipelines running, fixing deployment scripts, babysitting infrastructure — none of that is where the real value is.

So here’s the deal: don’t overcomplicate things. The simpler the system, the easier it is to find your way around it. Build what you need. The less you own, the more you can do — and the bill shrinks as a side effect. Time is money, and money is money, but navigability comes first: a system nobody understands is expensive no matter how cheap it looks on paper.

See also the Frugal Architect laws: https://thefrugalarchitect.com/laws/

Standardize

Standardizing is about being smart and efficient. Agility doesn’t come from reinventing the wheel every time. It comes from using proven, standardized processes that let you move quickly without getting bogged down in repetitive tasks. Standardize the way you log, test, build, deploy, and monitor your systems. When you use the same practices across the board, you reduce friction, simplify troubleshooting, and make it easier to scale.

But here’s the thing: standardization is not a goal, it’s a tool. It’s a means to help you move faster and more reliably. You’re not standardizing just for the sake of it — you’re doing it to create efficiency and consistency. But don’t forget the Share Nothing principle. Standardization doesn’t mean coupling your systems or teams in ways that restrict flexibility. Use standard practices to stay efficient, but don’t tie yourself down to anything that limits your independence or makes you less resilient. Standardize, but keep your systems loosely coupled and flexible.

How to use them

Principles by themselves are absolutes. And it’s almost impossible to satisfy all five of them to a 100% degree. And this is not the point. The point is to encourage smarter design, foster a constructive discussion, and provide vectors for negotiations. At the end of the day, designing software is all about tradeoffs. But to do a proper trade (or not), first, we need to understand what we are trading. This is what these principles are about — understanding the trade-offs of your system design and making decisions.

In my head, it looks like a radar chart, where each principle is represented on a 0 to 10 axis.

radar chart

The values are anecdotal — they can’t be measured like latency or error budgets, and they don’t describe the physical state of the system. They exist to make disagreements visible: if you score Share Nothing a 4 and I score it a 7, we finally know what to argue about.

But beware, the principles are not fighting against each other — they complement each other. The discussion should cover all five axes. Picking one principle and focusing on it alone will not result in better system design.
If we want to add more event-drivenness to our system, it will result in less sharing, and potentially better user experience, but will require more standardization and will increase costs. Raising the stakes on one principle will affect the other four.

Yes, we want everything to be 10/10, but that’s simply not possible.