Cell-Based Architecture: Capping Blast Radius With a Dumb Router and N Identical Stacks
Architecture · Advanced · 7 min read · published
This article was written by Claude (Anthropic) and published automatically.
What this solves: You have one big multi-tenant stack where a single bad deploy, poison record, or hot tenant takes down 100% of customers. This shows how to slice the whole stack into independent cells so any single failure is contained to a known fraction of traffic.
The Forces at Play
You start with one stack: one load balancer, one app tier, one database. It scales fine for a while. Then three pressures arrive at once.
Correlated failure. Every incident is a 100% incident. A poison message, a runaway query from one tenant, a bad index migration, an OOM caused by one customer's 400MB payload — all of it lands on every customer simultaneously. Your availability math is brutal: one shared component with three nines caps the whole product at three nines.
Vertical ceilings. The database is the usual one. You've sharded reads, added a cache, and you're still at 80% on a single primary that can't be split without a multi-quarter project.
Change risk. Deploys are the leading cause of outages, and a single stack means a deploy is a coin flip against your entire customer base at once.
The obvious answer is microservices — decompose vertically by domain. But that solves coupling between teams, not blast radius. Split into 30 services and a single bad deploy of the auth service still takes out 100% of users. Cell-based architecture makes the orthogonal cut: keep the stack whole, and replicate it horizontally into N independent copies, each serving a fixed slice of traffic. Failure stops being a boolean and becomes a fraction you choose.
The Shape
A cell is a complete, self-sufficient instance of your application — compute, data, cache, queues — sized to a maximum capacity and never talking to another cell. In front sits a deliberately thin router. Beside it, a control plane that owns placement but is not in the request path.
flowchart TB
C[Clients] --> R["Cell Router (thin)<br/>tenant_id → cell_id<br/>in-memory map, TTL 60s"]
R -->|"tenant A, C, F"| CELL1
R -->|"tenant B, D"| CELL2
R -->|"tenant E, G"| CELL3
subgraph CELL1 ["Cell 01 — us-east-1a/b"]
LB1[ALB] --> API1[API pods]
API1 --> DB1[(Postgres 01)]
API1 --> RD1[(Redis 01)]
API1 --> Q1[[SQS 01]]
Q1 --> W1[Workers 01]
W1 --> DB1
end
subgraph CELL2 ["Cell 02"]
LB2[ALB] --> API2[API pods] --> DB2[(Postgres 02)]
end
subgraph CELL3 ["Cell 03"]
LB3[ALB] --> API3[API pods] --> DB3[(Postgres 03)]
end
CP["Control Plane<br/>placement, provisioning,<br/>migration, deploy waves"]
CP -.->|"pushes routing table"| R
CP -.->|"creates / drains"| CELL1
CP -.->|"creates / drains"| CELL2
CP -.->|"creates / drains"| CELL3
OBS["Per-cell SLO dashboards<br/>cell_id on every metric"]
CELL1 -.-> OBS
CELL2 -.-> OBS
CELL3 -.-> OBS
style R fill:#fde68a,stroke:#b45309
style CP fill:#bfdbfe,stroke:#1e40af
The two things to notice: the dotted lines from the control plane never carry user requests, and there is no horizontal arrow between cells. If you find yourself drawing one, you don't have cells — you have a distributed monolith with extra hops.
How Data Flows Through It
A POST /v2/invoices from tenant acme:
- DNS resolves to the global router's anycast IP. The router terminates TLS and reads the tenant from the JWT claim (not the request body — routing must happen before parsing untrusted payloads).
- The router looks up
acmein its in-memory routing table:cell-02. This is a hash-map hit, sub-microsecond. If the tenant is missing, it calls the control plane's placement API synchronously once, caches the result, and proceeds. That's the only time the control plane is in a request path — and only for a tenant's first-ever request. - The router proxies to
cell-02's ALB with anX-Cell-Id: cell-02header. The cell rejects requests whose header doesn't match its own identity — a cheap invariant that catches routing bugs immediately instead of writing tenant data into the wrong shard. - Inside cell-02: API pod validates, writes to Postgres-02, enqueues a PDF render job on SQS-02. Nothing in this path can reach Postgres-01.
- Worker in cell-02 picks up the job, renders, writes back to Postgres-02, and pushes an object to an S3 prefix scoped to
cell-02/. - Metrics and logs emit with
cell_id=cell-02. The SLO dashboard shows 6 lines, not 1 — a cell going bad is visible as divergence, long before the aggregate moves.
Now the failure path: cell-02's Postgres primary fails over and takes 40 seconds. Tenants acme and dunder see errors. The other four cells don't notice, because there is no shared connection pool, no shared cache to stampede, no shared queue to back up. Your incident is titled "partial outage, 2 of 6 cells, 31% of tenants" and it stayed that way without anyone paging.
What Each Piece Owns
The router owns exactly one decision: request → cell. It must be the dumbest, most boring component you operate — no business logic, no authorization decisions, no schema knowledge, no database. It does not own retries across cells (retrying elsewhere would defeat isolation), and it does not own the source of truth for placement; it holds a cached projection.
The control plane owns placement, provisioning, cell draining, tenant migration, and deploy orchestration. It deliberately does not serve user traffic. If it's down, existing tenants keep working — you just can't onboard new ones or rebalance. That's the whole point of the split: control-plane availability requirements are an order of magnitude looser than data-plane ones.
A cell owns the complete lifecycle of its tenants' data and its own capacity ceiling. It does not own knowledge of how many cells exist, which cell it is relative to others, or any cross-tenant global state. A cell should be able to run in isolation with the router pointed straight at it — that's your test for whether it's truly self-sufficient.
Observability owns cross-cell aggregation. This is the one place a global view is legitimate, because it's read-only and out of band.
The contract between router and control plane stays tiny:
{
"version": 4127,
"generated_at": "2025-06-11T09:14:02Z",
"default_cell": "cell-03",
"assignments": [
{ "tenant_id": "acme", "cell_id": "cell-02", "state": "active" },
{ "tenant_id": "initech","cell_id": "cell-01", "state": "migrating",
"target_cell_id": "cell-04", "write_mode": "reject_with_503" }
]
}
The migrating state with an explicit write_mode is the part people forget. Tenant migration is a small, deliberate outage for one tenant, and encoding it in the routing table beats inventing dual-write machinery.
Where It Breaks Down
The first bottleneck is the thing you forgot to duplicate. In practice it's always one of: the global auth/user table, a shared Redis for rate limiting, a shared Kafka cluster, or a single Stripe webhook endpoint that fans out to cells. One shared synchronous dependency and your six-cell architecture has the availability of a one-cell architecture — but with six times the operational surface. Before you claim a blast-radius number, trace every outbound dependency of a cell and ask "if this is down, how many cells fail?"
Hot cells. Cells have a fixed ceiling, and tenants are not uniform. One enterprise customer growing 10x turns their cell into a permanent capacity emergency while five cells idle. You need per-cell utilization alarms and a rehearsed migration runbook, or you'll be doing your first tenant migration under pressure at 2am.
Deploys become the dominant operational burden. Six cells means six of everything: schema versions, config drift, certificate rotations, half-finished migrations. If you deploy all cells at once you've re-coupled them through your pipeline; if you deploy in waves you need automated health gating between waves or humans become the bottleneck. Cell architecture demands real deployment automation as a prerequisite, not a follow-up.
Partial failure gets subtler, not simpler. A cell that's 40% degraded — slow but not failing health checks — will keep receiving its slice of traffic indefinitely. You need an explicit "evacuate cell" lever (mark tenants for migration, or route to a spare cell after a cold data restore) and you need to have pulled it in a game day, because the first real use will be during an incident.
Cross-cell features fight the model. Global search, org-wide analytics, tenant-to-tenant sharing. Each of these wants to read across cells and each is a crack in the isolation story. The workable answer is an async, read-only, eventually-consistent aggregation pipeline that is explicitly allowed to be stale and is never in a write path.
When This Is Overkill
Below roughly a few thousand tenants and a database your largest instance type still handles comfortably, the correct design is one stack plus per-tenant limits: hard rate limits, query timeouts, connection quotas, and payload caps enforced per tenant. That's 90% of the noisy-neighbour protection for 5% of the operational cost. Add read replicas and a canary deploy stage before you add cells — a canary stage gives you real change-risk reduction with one extra environment, not N.
The intermediate step people skip: two cells. Prod-A and prod-B, 50/50. It forces you to eliminate shared dependencies and build wave deploys while the operational load is still trivial, and it already halves your blast radius.
The signals that you've genuinely outgrown a single stack:
- Your incident reviews keep concluding "one tenant's behaviour degraded everyone," and you've already added quotas.
- The primary database is approaching a hard vertical ceiling and a functional split won't help because the load is one dominant table.
- A customer contract or regulator requires isolation or data residency you can't satisfy with row-level scoping.
- Availability targets have moved to four nines and your math says a single shared data tier cannot get there.
If none of those are true, cells will cost you six sets of pages and buy you a nicer architecture diagram.
Key takeaway: A cell only limits blast radius if nothing in the request path is shared across cells — audit for the one global database, cache, or coordinated deploy that silently re-couples them.
Real-world challenge
Your team runs 6 cells and reports "blast radius = 16%" to leadership. During a routine release, error rates spike to 90% globally for 11 minutes. Post-incident, the app image had a bug that crash-looped on startup when an env var was missing. The cells are genuinely independent at the data layer — separate DBs, separate caches, separate queues. Why did all 6 fail, and what do you change?
Diagnosis
The data plane was isolated; the control plane was not. Check your deploy pipeline: if the release job runs for cell in cells; do deploy $cell; done with no gate between iterations, or worse deploys them in parallel, then your deployment system is a shared dependency that touches every cell within seconds. Cell isolation is only as good as the slowest thing that can change all cells at once.
Confirm it by looking at the timestamps of the first crash-loop in each cell. If they're within a minute of each other, it's the pipeline, not the code.
Fix
Make deploys a wave-based, health-gated progression, and treat the cell list as ordered:
waves:
- cells: [cell-canary] # internal tenants only
bake: 30m
- cells: [cell-01]
bake: 1h
- cells: [cell-02, cell-03]
bake: 1h
- cells: [cell-04, cell-05]
gate:
promote_if:
error_rate_5xx: "< 0.5%"
p99_latency_ms: "< 400"
on_breach: halt_and_rollback_wave
Then add the guardrails that make the gate real:
- No global apply. Remove any credential or job that can write to all cells in one action; require one pipeline run per wave.
- Startup validation. Fail config validation in CI, not at container start — a missing env var should never be discoverable only in production.
- Rehearse the claim. Quarterly, deliberately break one cell in prod and verify the observed impact matches the number you promised leadership. An untested blast-radius number is a guess.