Feature Flag Check Adds Latency to Every Request: Move Evaluation In-Process
Feature Flags · Intermediate · 7 min read · published
This article was written by Claude (Anthropic) and published automatically.
What this solves: If every flag lookup makes a network call, your p99 inherits the flag vendor's p99. Here's the delivery-plane/evaluation-plane split that removes that hop.
The Forces at Play
Someone opens a flame graph and finds that a feature flag check adds latency to every request — 4ms here, 9ms there, and one endpoint that checks eleven flags now carries 40ms of pure flag lookup. The obvious implementation (GET /evaluate?flag=x&user=y) is correct, auditable, and instantly consistent. It is also a synchronous network call to a third party sitting on your critical path.
The tensions:
- Latency vs. freshness. A kill switch should take effect in seconds. But "seconds" does not require a round trip per evaluation.
- Availability coupling. Remote evaluation means your checkout availability is now
your_uptime × flag_vendor_uptime. Flags are used most during incidents, which is exactly when you don't want a new dependency. - Privacy. Remote evaluation means shipping user attributes (email, plan, country) to a vendor on every request.
- Correctness of analytics. If nothing calls home, how do you know who saw which variant?
The pattern that resolves this splits the system into a delivery plane (push rules out) and an evaluation plane (apply them locally), and pushes analytics onto a separate asynchronous path.
The Shape
flowchart TB
subgraph CP["Control plane (writes, low volume)"]
UI[Admin UI / API] --> FS[Flag service]
FS --> DB[(Flag config DB<br/>+ audit log)]
FS --> COMP[Ruleset compiler<br/>config -> immutable snapshot v42]
COMP --> STORE[(Snapshot store)]
STORE --> CDN[[CDN / edge cache<br/>GET /rules/env/latest]]
COMP --> BUS[Change stream<br/>SSE / WebSocket]
end
subgraph DP["Data plane (reads, very high volume)"]
SDK[SDK in app process]
MEM[[In-memory ruleset v42]]
EVAL{{"evaluate(flag, context)<br/>pure function, ~200ns"}}
BUF[Exposure buffer<br/>batched, dropped-on-full]
end
CDN -- "bootstrap fetch, once at boot" --> SDK
BUS -- "invalidation ping (no payload)" --> SDK
SDK --> MEM
APP[Request handler] --> EVAL
MEM --> EVAL
EVAL --> APP
EVAL -.-> BUF
BUF -- "flush every 5s / on shutdown" --> ING[Exposure ingest]
ING --> WH[(Warehouse / experiment analysis)]
SNAP[Baked-in snapshot in image] -. "synchronous fallback at boot" .-> MEM
The crucial line is the one that isn't there: no arrow from EVAL back to the control plane. The request path touches only process memory.
How Data Flows Through It
An engineer ramps new-checkout from 5% to 50% for plan=pro in us-east:
- Write. Admin UI → flag service → config DB, with an audit row. This is a handful of writes per day, so it can be a boring Postgres row.
- Compile. The ruleset compiler materialises all flags for that environment into one immutable, versioned snapshot. Compilation resolves segment references (
segment:power_users→ its predicate tree) so the SDK never has to chase pointers. Output is content-addressed:rules/prod/v42.json, plusrules/prod/latestpointing at it. - Distribute. The snapshot is pushed to a CDN with a long max-age on the versioned URL and a short one on
latest. A change-stream connection tells connected SDKs "there's a new version" — a ping, not the payload, so the CDN absorbs the thundering herd. - Refresh. Each SDK fetches the new snapshot and atomically swaps its in-memory pointer. Old evaluations in flight finish against v41; new ones see v42. No lock, no partial state.
- Evaluate. A request arrives. The handler calls
flags.variant("new-checkout", { userId, plan, country }). The SDK walks the rule list, finds theplan=pro AND country=USrule, then does a deterministic bucket:murmur3(flagKey + ":" + salt + ":" + userId) % 10000 < 5000. Same user, same flag, same answer on every pod — no coordination needed. - Record. The evaluation appends
{flag, variant, userId, rulesetVersion, ts}to a bounded in-memory buffer. Flushed every few seconds, on buffer-full, and on SIGTERM. If the ingest endpoint is down, exposures are dropped — deliberately, because analytics must never back-pressure checkout.
End-to-end propagation: typically under two seconds. Request-path cost: a hash and a few comparisons.
What Each Piece Owns
Flag service + config DB own the intent: who changed what, when, and why. It does not own evaluation, and it must never be in the read path.
Ruleset compiler owns flattening config into a self-contained, versioned artifact. It does not know about users — it never sees a user ID. That's what keeps it cheap and cacheable.
CDN owns fan-out to thousands of processes. It does not own correctness; a stale CDN just means a slightly older ruleset version, which is a bounded, acceptable failure.
Change stream owns freshness, and only freshness. It carries no payload, so if it's down, SDKs degrade to polling every 30–60s. Losing it slows propagation; it never breaks evaluation.
SDK owns the in-memory snapshot, the bucketing hash, and the exposure buffer. It deliberately does not own attribute lookup — it cannot call your user service. Every attribute must be handed in by the caller:
interface EvaluationContext {
key: string; // bucketing unit: user, account, or device
kind: 'user' | 'account';
attributes: Record<string, string | number | boolean>;
}
interface Ruleset {
version: number; // monotonic; log this with every exposure
environment: string;
flags: Record<string, {
defaultVariant: string;
salt: string; // stable per flag: changing it reshuffles buckets
rules: Array<{
when: Predicate; // pre-resolved, no segment references
rollout: Array<{ variant: string; weightBps: number }>; // sums to 10000
}>;
}>;
}
That version field is what turns "it worked on my pod" into a diagnosable event.
Where It Breaks Down
Cold start is the first real bug. Between process start and the first successful fetch, the SDK has no ruleset and returns the code default. Deploy 200 pods and a fraction of traffic briefly gets the wrong variant — and it correlates with deploys, so it looks like a deploy bug. Fix: bake a snapshot into the image, load it synchronously, gate readiness on initialization.
Ruleset size is the scaling wall. Local evaluation means every process holds every flag. At 4,000 flags with big userId IN (...) allowlists, you're shipping multi-megabyte JSON to every pod on every change, and Lambda cold starts start to hurt. Symptoms: memory growth proportional to flag count, refresh CPU spikes. Mitigations: per-service ruleset scoping, archiving stale flags aggressively, replacing ID lists with hashed segments.
Split-brain during rollout. Pods refresh at slightly different times, so for a second or two some serve v41 and some v42. For a boolean ramp that's fine. For two flags that must change together (new API + new client behaviour) it is not — a user can hit an old pod then a new one. Fix: one flag with multiple variants, or a snapshot version pinned per session.
Exposure loss is silent. Buffers die with the pod. You notice weeks later when an experiment's sample count doesn't match page views. Always flush on SIGTERM; alert on exposures-per-minute dropping.
Operational burden lands on flag hygiene, not infrastructure. The delivery plane is nearly free to run. What actually costs you is 3,000 permanent flags nobody dares delete, each an untested branch. Enforce expiry dates and fail CI on flags older than 90 days.
When This Is Overkill
If you have one service, a few dozen flags, and no experimentation programme, the correct architecture is a config file or a table read into memory at boot, refreshed on a 30-second timer:
setInterval(async () => { rules = await loadRulesFromDB(); }, 30_000);
export const isEnabled = (k: string, userId: string) =>
rules[k]?.enabled && hash(k + userId) % 100 < (rules[k].pct ?? 100);
That's ~30 lines, zero vendors, and it already gives you the two things that matter most: no network call on the request path, and consistent bucketing.
The specific signals you've outgrown it:
- More than one service needs the same flag and they disagree about its state.
- You need a kill switch to propagate in seconds, and 30-second polling against your primary DB is now a load concern.
- Someone asks "which users saw variant B last Tuesday?" — you now need exposure events, which means the asynchronous analytics path.
- Non-engineers need to flip flags, which means an audit log and permissions, which means a control plane.
Until then, the distributed version buys you complexity you're not yet being paid for.
Key takeaway: Ship the rules to the process, not the question to the server: flag evaluation should be a pure function over an in-memory ruleset, with staleness measured in seconds, not a blocking RPC on the request path.
Real-world challenge
You migrated from remote flag evaluation to a local-eval SDK. p99 dropped nicely, but now two things happen: (1) during deploys, a handful of pods briefly serve the control (off) variant for a flag that is 100% on, and (2) your experiment analysis shows ~3% fewer exposure events than page views for the same flag. Diagnose both.
Symptom 1 — control leak at startup. The SDK returns the code default until its first ruleset fetch completes. If your code calls isEnabled() before waitForInitialization(), or the CDN fetch is slow, you evaluate against an empty ruleset and silently get the default.
Fixes, in order of robustness:
- Bake a ruleset snapshot into the container image and load it synchronously at boot, then refresh over the network.
- Block readiness (not liveness) until initialization completes.
- Make the code default equal the current production value, and alert when they diverge.
await flags.waitForInitialization({ timeoutMs: 500 });
// falls back to the baked-in snapshot, never to an empty ruleset
app.listen(PORT);
Symptom 2 — missing exposures. Exposure events are batched in memory and flushed on a timer. Pods that get SIGTERM'd mid-interval drop their buffer, which matches a small, deploy-correlated shortfall. Add a flush on shutdown and give the pod a preStop grace period:
process.on('SIGTERM', async () => { await flags.flushEvents(); server.close(); });
If exposures matter for billing or experiment validity, also persist the buffer to disk or accept the loss explicitly with a documented error bar.