Feature Flag Check Adds Latency to Every Request: Move Evaluation In-Process

Feature Flags · Intermediate · 7 min read · published

This article was written by Claude (Anthropic) and published automatically.

What this solves: If every flag lookup makes a network call, your p99 inherits the flag vendor's p99. Here's the delivery-plane/evaluation-plane split that removes that hop.

The Forces at Play

Someone opens a flame graph and finds that a feature flag check adds latency to every request — 4ms here, 9ms there, and one endpoint that checks eleven flags now carries 40ms of pure flag lookup. The obvious implementation (GET /evaluate?flag=x&user=y) is correct, auditable, and instantly consistent. It is also a synchronous network call to a third party sitting on your critical path.

The tensions:

The pattern that resolves this splits the system into a delivery plane (push rules out) and an evaluation plane (apply them locally), and pushes analytics onto a separate asynchronous path.

The Shape

flowchart TB
  subgraph CP["Control plane (writes, low volume)"]
    UI[Admin UI / API] --> FS[Flag service]
    FS --> DB[(Flag config DB<br/>+ audit log)]
    FS --> COMP[Ruleset compiler<br/>config -> immutable snapshot v42]
    COMP --> STORE[(Snapshot store)]
    STORE --> CDN[[CDN / edge cache<br/>GET /rules/env/latest]]
    COMP --> BUS[Change stream<br/>SSE / WebSocket]
  end

  subgraph DP["Data plane (reads, very high volume)"]
    SDK[SDK in app process]
    MEM[[In-memory ruleset v42]]
    EVAL{{"evaluate(flag, context)<br/>pure function, ~200ns"}}
    BUF[Exposure buffer<br/>batched, dropped-on-full]
  end

  CDN -- "bootstrap fetch, once at boot" --> SDK
  BUS -- "invalidation ping (no payload)" --> SDK
  SDK --> MEM
  APP[Request handler] --> EVAL
  MEM --> EVAL
  EVAL --> APP
  EVAL -.-> BUF
  BUF -- "flush every 5s / on shutdown" --> ING[Exposure ingest]
  ING --> WH[(Warehouse / experiment analysis)]

  SNAP[Baked-in snapshot in image] -. "synchronous fallback at boot" .-> MEM

The crucial line is the one that isn't there: no arrow from EVAL back to the control plane. The request path touches only process memory.

How Data Flows Through It

An engineer ramps new-checkout from 5% to 50% for plan=pro in us-east:

  1. Write. Admin UI → flag service → config DB, with an audit row. This is a handful of writes per day, so it can be a boring Postgres row.
  2. Compile. The ruleset compiler materialises all flags for that environment into one immutable, versioned snapshot. Compilation resolves segment references (segment:power_users → its predicate tree) so the SDK never has to chase pointers. Output is content-addressed: rules/prod/v42.json, plus rules/prod/latest pointing at it.
  3. Distribute. The snapshot is pushed to a CDN with a long max-age on the versioned URL and a short one on latest. A change-stream connection tells connected SDKs "there's a new version" — a ping, not the payload, so the CDN absorbs the thundering herd.
  4. Refresh. Each SDK fetches the new snapshot and atomically swaps its in-memory pointer. Old evaluations in flight finish against v41; new ones see v42. No lock, no partial state.
  5. Evaluate. A request arrives. The handler calls flags.variant("new-checkout", { userId, plan, country }). The SDK walks the rule list, finds the plan=pro AND country=US rule, then does a deterministic bucket: murmur3(flagKey + ":" + salt + ":" + userId) % 10000 < 5000. Same user, same flag, same answer on every pod — no coordination needed.
  6. Record. The evaluation appends {flag, variant, userId, rulesetVersion, ts} to a bounded in-memory buffer. Flushed every few seconds, on buffer-full, and on SIGTERM. If the ingest endpoint is down, exposures are dropped — deliberately, because analytics must never back-pressure checkout.

End-to-end propagation: typically under two seconds. Request-path cost: a hash and a few comparisons.

What Each Piece Owns

Flag service + config DB own the intent: who changed what, when, and why. It does not own evaluation, and it must never be in the read path.

Ruleset compiler owns flattening config into a self-contained, versioned artifact. It does not know about users — it never sees a user ID. That's what keeps it cheap and cacheable.

CDN owns fan-out to thousands of processes. It does not own correctness; a stale CDN just means a slightly older ruleset version, which is a bounded, acceptable failure.

Change stream owns freshness, and only freshness. It carries no payload, so if it's down, SDKs degrade to polling every 30–60s. Losing it slows propagation; it never breaks evaluation.

SDK owns the in-memory snapshot, the bucketing hash, and the exposure buffer. It deliberately does not own attribute lookup — it cannot call your user service. Every attribute must be handed in by the caller:

interface EvaluationContext {
  key: string;                 // bucketing unit: user, account, or device
  kind: 'user' | 'account';
  attributes: Record<string, string | number | boolean>;
}

interface Ruleset {
  version: number;             // monotonic; log this with every exposure
  environment: string;
  flags: Record<string, {
    defaultVariant: string;
    salt: string;              // stable per flag: changing it reshuffles buckets
    rules: Array<{
      when: Predicate;         // pre-resolved, no segment references
      rollout: Array<{ variant: string; weightBps: number }>; // sums to 10000
    }>;
  }>;
}

That version field is what turns "it worked on my pod" into a diagnosable event.

Where It Breaks Down

Cold start is the first real bug. Between process start and the first successful fetch, the SDK has no ruleset and returns the code default. Deploy 200 pods and a fraction of traffic briefly gets the wrong variant — and it correlates with deploys, so it looks like a deploy bug. Fix: bake a snapshot into the image, load it synchronously, gate readiness on initialization.

Ruleset size is the scaling wall. Local evaluation means every process holds every flag. At 4,000 flags with big userId IN (...) allowlists, you're shipping multi-megabyte JSON to every pod on every change, and Lambda cold starts start to hurt. Symptoms: memory growth proportional to flag count, refresh CPU spikes. Mitigations: per-service ruleset scoping, archiving stale flags aggressively, replacing ID lists with hashed segments.

Split-brain during rollout. Pods refresh at slightly different times, so for a second or two some serve v41 and some v42. For a boolean ramp that's fine. For two flags that must change together (new API + new client behaviour) it is not — a user can hit an old pod then a new one. Fix: one flag with multiple variants, or a snapshot version pinned per session.

Exposure loss is silent. Buffers die with the pod. You notice weeks later when an experiment's sample count doesn't match page views. Always flush on SIGTERM; alert on exposures-per-minute dropping.

Operational burden lands on flag hygiene, not infrastructure. The delivery plane is nearly free to run. What actually costs you is 3,000 permanent flags nobody dares delete, each an untested branch. Enforce expiry dates and fail CI on flags older than 90 days.

When This Is Overkill

If you have one service, a few dozen flags, and no experimentation programme, the correct architecture is a config file or a table read into memory at boot, refreshed on a 30-second timer:

setInterval(async () => { rules = await loadRulesFromDB(); }, 30_000);
export const isEnabled = (k: string, userId: string) =>
  rules[k]?.enabled && hash(k + userId) % 100 < (rules[k].pct ?? 100);

That's ~30 lines, zero vendors, and it already gives you the two things that matter most: no network call on the request path, and consistent bucketing.

The specific signals you've outgrown it:

Until then, the distributed version buys you complexity you're not yet being paid for.

Key takeaway: Ship the rules to the process, not the question to the server: flag evaluation should be a pure function over an in-memory ruleset, with staleness measured in seconds, not a blocking RPC on the request path.

Real-world challenge

You migrated from remote flag evaluation to a local-eval SDK. p99 dropped nicely, but now two things happen: (1) during deploys, a handful of pods briefly serve the control (off) variant for a flag that is 100% on, and (2) your experiment analysis shows ~3% fewer exposure events than page views for the same flag. Diagnose both.

Symptom 1 — control leak at startup. The SDK returns the code default until its first ruleset fetch completes. If your code calls isEnabled() before waitForInitialization(), or the CDN fetch is slow, you evaluate against an empty ruleset and silently get the default.

Fixes, in order of robustness:

  1. Bake a ruleset snapshot into the container image and load it synchronously at boot, then refresh over the network.
  2. Block readiness (not liveness) until initialization completes.
  3. Make the code default equal the current production value, and alert when they diverge.
await flags.waitForInitialization({ timeoutMs: 500 });
// falls back to the baked-in snapshot, never to an empty ruleset
app.listen(PORT);

Symptom 2 — missing exposures. Exposure events are batched in memory and flushed on a timer. Pods that get SIGTERM'd mid-interval drop their buffer, which matches a small, deploy-correlated shortfall. Add a flush on shutdown and give the pod a preStop grace period:

process.on('SIGTERM', async () => { await flags.flushEvents(); server.close(); });

If exposures matter for billing or experiment validity, also persist the buffer to disk or accept the loss explicitly with a documented error bar.