← blog · August 4, 2026

What an active-active Kubernetes architecture really takes

Active-active is not about splitting traffic, it is about writing consistently in two regions at once, and the real work ends up in the data layer. The cost of synchronous replication, why two regions cannot form a quorum, conflict resolution, session and cache decisions, which workloads should never be active-active, and a staged road map.

When a team says "let's go active-active", what they usually mean is: when a region dies, nobody should notice. That wish does not require active-active. Active-active means both regions accept write traffic at the same time, and that is a data modelling problem, not an infrastructure cloning problem. Splitting traffic across two regions is a few days of work. Writing consistently in two places at once is a redesign of the application.

What actually separates active-passive from active-active

In an active-passive setup the second region is running and receiving data, but it does not accept writes. On failure the primary role moves there. You measure two numbers: RTO, how long until you are serving again, and RPO, how much data you accepted losing. A well built active-passive setup that is drilled regularly gives you an RTO in minutes and an RPO in seconds.

Active-active drives both toward zero on paper. The price is that you no longer have a single authoritative copy of the data. Two requests can modify the same record in two regions at the same instant, and one of them has to win. That decision is either made by the database, and you pay in latency, or made by the application, and you pay in complexity, or made by nobody, and you lose data quietly. The third case is the most common, because it never raises an error anywhere.

The cost is not double either. Infrastructure is double, but operations are much worse: version parity across two clusters, secret rotation in two clusters, schema migrations in two clusters, double the observability surface, and the hardest one, people who can diagnose a two-region incident at three in the morning.

The traffic layer is the easy part

DNS based routing is the cheapest option and the slowest to fail over. Dropping your TTL to 30 seconds does not give you a 30 second failover. Some resolvers ignore TTLs, browsers and various runtimes keep their own caches, and corporate networks will serve a stale record for minutes. Use DNS for planned failover and coarse geographic steering, not for a reaction measured in seconds.

Anycast announces one IP from several places, and once routing tables converge, traffic drifts to whichever region is still up. The part people skip is that open TCP connections can land on a different node during convergence and get reset. That is invisible for short HTTP requests and very visible for long lived WebSocket connections.

The arrangement that works best in practice is anycast at the edge in front of a layer seven router that picks a region based on health. The mistake made most often here is a shallow health check. If the check is an endpoint that returns 200 unconditionally, a region that cannot write to its database still looks healthy and keeps taking traffic. The check has to prove the region can do real work, which means it has to touch the write path.

On the Kubernetes side the call is clear: one cluster per region. Stretching a single cluster across two regions turns etcd into a cross-region Raft group. The etcd tuning guide says the heartbeat interval should be close to the average round trip time between members, roughly 0.5 to 1.5 times it, and the election timeout should be at least 5 to 10 times the heartbeat. The 50 second upper bound for the election timeout that the guide documents is explicitly reserved for globally distributed clusters. So it is possible, but you have pushed your control plane's leader election into the tens of seconds and every bit of network jitter triggers a new one. Separate clusters plus a routing layer above them produce far fewer surprises.

Inside a cluster, the trafficDistribution field that went GA in Kubernetes 1.33 is worth using. In 1.34 a clearer alias, PreferSameZone, was added for PreferClose:

apiVersion: v1
kind: Service
metadata:
  name: api
spec:
  trafficDistribution: PreferClose
  selector:
    app: api
  ports:
    - port: 80
      targetPort: 8080

This is zone level, not region level, so it is not a multi-region answer. It is the same principle at small scale, which makes it a good first step for a team that has not yet felt what locality does to latency numbers.

The data layer is where the work is

You cannot argue with physics. Light travels through fibre at roughly 200,000 km per second, so every 100 km costs about 1 ms of round trip. Two regions 1000 km apart give you 10 ms in the best case and 15 to 40 ms with real routing. Under synchronous replication that time is added to every commit.

synchronous_standby_names = 'ANY 1 (replica_a, replica_b)'
synchronous_commit = remote_apply

remote_apply waits until the standby confirms the transaction has been applied to its database. remote_write only waits for confirmation that the record was written out, which is faster but leaves a data loss window if the standby's operating system crashes. Choose remote_apply across regions and a workload doing 200 small writes per second may drop to 25. That is not a bug, it is the bill. The problem is that nobody measures the bill until it arrives in production.

Then comes quorum: two regions cannot form a majority. When a system splits in half, both halves conclude the other one died. Every quorum based component, etcd included, needs a third failure domain. That third site does not need a full replica, a small voting witness is enough. If you do not have a third failure domain, do not enable automatic failover; gate it behind a human. Two regions plus automatic failover is the shortest path to split-brain.

For conflict resolution there are three options and none of them are pleasant:

  • Last write wins. Easy to implement, loses data silently. With clock skew, the "last" writer may not be the last writer.
  • CRDTs. They genuinely solve a narrow class of problems: counters, sets, collaborative text. They do not generalise to business rules like orders, payments or stock levels.
  • Record ownership. Each record belongs to a region, writes for it are routed there, the other region only reads. This distributes the complexity most honestly and it is the only approach most teams can actually implement. The ownership key is usually a customer or tenant id.

The first concrete problem you hit on the way to record ownership is auto-incrementing primary keys. Both regions start counting at one and collide the moment you merge anything. The fix is either collision-free identifiers or per-region sequences with distinct start and step values. The second is simple until the day you add a third region.

Sessions and caches

Sticky sessions and active-active eat each other. If a session lives in one region's memory, a user landing in the other region loses it. If you try to keep sessions synchronised across regions, every request pays a cross-region round trip. The practical answer is to make the session portable: a signed, short lived token that any region can validate locally. The price is revocation. You have to distribute a revocation list, but a small list propagated once a second is far cheaper than a cross-region lookup on every request.

For caches the rule is simple: each region gets its own cache, and you never build a shared cache cluster spanning regions. Invalidation messages arrive at least once and out of order across regions. Version your keys instead of deleting them, because order dependent invalidation produces different results in the two regions and is miserable to debug.

Latency reshapes application design

A query that takes half a millisecond inside a region takes 40 ms across regions. A classic N+1 endpoint doing a hundred small queries costs 50 ms locally and 4 seconds when the data lives in the other region. This is usually the first shock for teams moving to active-active, and it is not an architectural fault; a code pattern that was tolerated for years just had its bill called in.

The second shock is losing read-your-own-writes. The user writes to one region, the next request lands in the other, they do not see their change, and they retry. Either you pin the session to a region for its lifetime, which means giving up part of active-active, or you route reads back to the writing region for a short window after each write.

The third is retries. When a timed-out request is retried in the other region, the first attempt may well have succeeded. In an active-active system idempotency keys are not a nicety on the payment endpoint, they are mandatory on every write path.

What should not be active-active

  • Counters that need strong consistency: inventory reservations, quotas, balances, seat allocation. Keep a single writer.
  • Scheduled work. The same cron sitting in two regions runs twice. Reconciliation jobs, invoice generation, outbound email and batch notifications all belong here. Give them a lease mechanism or leave them in one region.
  • Schema migrations. Two schema versions writing the same data concurrently produce errors you cannot undo. Under active-active, migrations become backwards compatible and multi-step by necessity.
  • Single-writer data stores and products that were never designed for multi-primary operation. If the documentation does not use that word, do not pretend otherwise.

A staged road map

  1. Make one region reproducible. Cluster build, networking, secrets and certificates defined as code. The day you restore a backup into an empty environment and watch it work is your real starting point.
  2. Build active-passive and drill it. Continuous replication, failover automation, a measured RTO. A standby region that is never exercised is an expensive hope. For most teams, everything they actually needed is contained in this step.
  3. Serve reads from both. Read replicas plus a clear marking of which endpoints tolerate stale reads. The win is double: latency drops and the second region is continuously proven.
  4. Open a narrow write path. Use record ownership, and start somewhere low risk: user preferences, profiles, telemetry.
  5. Open general writes only if the data store was designed for it.

Deciding never to go past step three is a perfectly good engineering decision. Stopping deliberately beats a half-finished active-active setup by a wide margin.

The question worth asking

Which failures actually hurt you? Full region outages are rare. Bad releases, broken migrations, full disks, expired certificates and misconfiguration cause far more downtime. Active-active fixes none of those, and it makes most of them worse, because a bad release ships to both regions at once. The single thing it solves is losing an entire region.

Make the investment decision with that in mind. List your outages from the past year and write next to each one whether active-active would have prevented it. For most teams that column comes back empty, and an empty column says the budget for the second region belongs in observability, migration safety, and an active-passive setup that somebody actually rehearses.