TL;DR: In a Node.js API, put the Express feature flag check in middleware after authentication, return 404 or 403 before expensive route work starts, and attach one immutable decision record to the request. For an edtech AI experiment, that record should contain a tenant pseudonym, cohort, flag revision, decision, and cost bucket. The handler can then report model usage against the exact decision that admitted the request, which makes cohort comparisons reproducible without copying student data into telemetry.
The least complex design is one gate and one context object. Do not let each lesson-generation handler fetch flags, choose a cohort, and invent its own metric labels. A request should move through authentication, tenant context, feature decision, application work, and usage recording in that order. In Express terms, the gate has the familiar req, res, next responsibility; the example below uses Python because the important contract is independent of a web framework.
How should Express middleware check a feature flag before API route gating?
The middleware should answer a narrow question: may this authenticated tenant invoke this route under the current flag revision? It should not decide whether the generated lesson is good, calculate a bill, or inspect raw student content. Those belong to evaluation, accounting, and application layers.
For a cohort experiment, the input can stay small: the flag key, tenant ID, entitlement, rollout percentage, and configuration revision. The output is smaller still: allow or deny, plus the cohort and reason. Martin Fowler's feature-toggle taxonomy separates release, experiment, ops, and permissioning concerns; keeping the decision explicit matters because those categories have different lifetimes and owners. Here, entitlement is a hard boundary, while cohort assignment is an experiment mechanism.
Order matters. Authentication must establish the tenant before the gate runs. The gate must run before retrieval, prompt construction, or a model call, because a denied request should consume none of those resources. The usage recorder runs after admitted work and uses the decision already attached to the request. It does not evaluate the flag again.
That last detail prevents a subtle accounting error. If configuration changes during a request, a second evaluation can label usage with a different revision or cohort than the one that actually opened the route. One request, one decision.
Consider one concrete request from tenant-042. Authentication resolves the tenant and its entitlement, then the middleware reads revision rev-17, assigns the tenant to one stable bucket from 0 through 99, and stores the resulting decision. If the tenant lands in control, the request ends before retrieval and its decision event records zero model usage; if it lands in treatment, the handler runs and the model adapter reports usage into the decision's cost bucket. Suppose an operator raises the rollout while that request is active. The handler still uses rev-17 from request context rather than asking the evaluator again. This is the trade-off I would choose: a decision can be a few seconds older than the control plane, but the accounting record stays internally consistent. A later request sees the newer revision and gets attributed separately. Mixing those two revisions would make a cohort total look precise while quietly combining different populations. That is a worse failure than a short propagation delay because an eval-driven release depends on being able to replay which rule admitted which work.
Build the gate and cost ledger
This runnable example models the middleware contract without coupling it to a flag service or web framework. The rollout uses a deterministic SHA-256 bucket, so the same tenant and experiment key remain in the same cohort for a given rule. The percentages and token counts are illustrative policy inputs, not performance claims or vendor prices.
from __future__ import annotations
from dataclasses import dataclass, field
from decimal import Decimal
import hashlib
from typing import Callable
@dataclass(frozen=True)
class FlagRule:
key: str
revision: str
enabled: bool
rollout_percent: int
@dataclass(frozen=True)
class FeatureDecision:
flag_key: str
revision: str
cohort: str
allowed: bool
reason: str
@dataclass
class Request:
tenant_id: str
entitled: bool
context: dict[str, object] = field(default_factory=dict)
@dataclass(frozen=True)
class Response:
status: int
body: dict[str, object]
def stable_bucket(tenant_id: str, flag_key: str) -> int:
material = f"{flag_key}:{tenant_id}".encode("utf-8")
digest = hashlib.sha256(material).digest()
return int.from_bytes(digest[:8], "big") % 100
def decide(request: Request, rule: FlagRule) -> FeatureDecision:
if not request.entitled:
return FeatureDecision(
rule.key, rule.revision, "ineligible", False, "not_entitled"
)
if not rule.enabled:
return FeatureDecision(
rule.key, rule.revision, "eligible", False, "flag_disabled"
)
in_experiment = stable_bucket(request.tenant_id, rule.key) < rule.rollout_percent
cohort = "treatment" if in_experiment else "control"
return FeatureDecision(
rule.key, rule.revision, cohort, in_experiment, "cohort_assignment"
)
Handler = Callable[[Request], Response]
def feature_gate(rule: FlagRule, handler: Handler) -> Handler:
def middleware(request: Request) -> Response:
decision = decide(request, rule)
request.context["feature_decision"] = decision
if not decision.allowed:
return Response(404, {"error": "not_found"})
return handler(request)
return middleware
def generate_lesson(request: Request) -> Response:
decision = request.context["feature_decision"]
assert isinstance(decision, FeatureDecision)
# These values stand in for usage returned by the model adapter.
input_tokens = 820
output_tokens = 240
request.context["usage"] = {
"input_tokens": input_tokens,
"output_tokens": output_tokens,
"cost_units": Decimal("0.001060"),
"cost_bucket": f"{decision.flag_key}:{decision.revision}:{decision.cohort}",
}
return Response(200, {"lesson_status": "ready"})
rule = FlagRule(
key="ai_lesson_outline",
revision="rev-17",
enabled=True,
rollout_percent=25,
)
route = feature_gate(rule, generate_lesson)
example = route(Request(tenant_id="tenant-042", entitled=True))
print(example.status)
In an Express application, feature_gate maps to a middleware factory. Store the immutable FeatureDecision on res.locals or another request-scoped context, send the denial response immediately, and call next() only for an allowed request. Keep the underlying evaluator behind a small interface. The route should not care whether rules come from a local file, a database, or a remote control plane.
There is a deliberate trade-off in the sample: the rollout unit is the tenant, not the student. A school district should not have teachers in treatment while other teachers from the same tenant unexpectedly see control behavior when the experiment is meant to compare tenant cohorts. A student-level experiment may be valid for a different research question, but it changes consent, analysis, and privacy requirements. Choose the unit before choosing the hash key.
The 404 response is also a policy choice. It avoids confirming the existence of an unavailable capability. An API whose clients need to distinguish missing entitlement from a missing route can return 403 instead, provided that contract is documented and tested. The important behavior is terminal: denied requests do not fall through to the handler.
Observe decisions without collecting lesson content
The useful join key is the decision, not the prompt. Emit a decision event when the gate runs and a usage event after admitted work finishes. Both should carry the same request correlation ID and low-cardinality fields such as flag key, revision, cohort, decision reason, and a pseudonymous tenant key. Usage can include input tokens, output tokens, latency, and internally calculated cost units.
Do not put student names, lesson text, prompt bodies, or raw tenant IDs into flag telemetry. GDPR Article 5 includes data minimization: personal data should be adequate, relevant, and limited to what is necessary for the purpose. A cohort cost comparison needs stable attribution, but it does not need the content being taught.
Cardinality deserves attention too. Tenant-level analysis requires a tenant dimension somewhere, yet putting every tenant ID into every metrics label can make a time-series system awkward and expensive. A practical split is aggregated counters for dashboards and structured decision/usage events for tenant-level analysis. Keep the pseudonym mapping in a controlled data system, outside the event payload.
Now the experiment query has a defensible denominator. Compare admitted requests, token usage, and cost units by cohort and revision. Keep evaluation quality beside cost rather than collapsing both into one number: a cheaper cohort that produces worse lesson outputs has not won. The eval harness should use the same revision and cohort fields, allowing reviewers to compare quality, spend, and failure rate over the same population.
No prompt logging required.
Failure behavior is part of the feature
A gate is production control flow, so ambiguous states need explicit policy. If the rule is absent, malformed, or outside its validity window, a permission-like feature should fail closed. If an experiment evaluator is temporarily unavailable, use a previously validated snapshot only when the team has defined its age limit and rollback semantics. Otherwise deny the experimental path and keep the established path available.
Retries must not reshuffle a tenant. Deterministic assignment handles that as long as the hash inputs and algorithm remain stable. Changing the flag key, bucketing algorithm, or rollout unit creates a new population; treat that as a new revision and do not merge its measurements casually with the old one.
Test the boundary, not just the happy path. A compact suite should cover an unauthenticated request never reaching the gate, an unentitled tenant receiving the chosen denial status, disabled and missing rules, stable assignment across retries, exact bucket boundaries at 0 and 100 percent, and a handler that records usage under the attached revision. Add a concurrency test that changes the rule source after admission and confirms that the handler still reports the original decision.
There is another easy trap: middleware registration order can make a correct gate useless. Registering the route handler first means expensive work may already have started. An integration test should use a handler spy and prove it has zero calls for every denied case.
Ship the experiment as an accounting contract
Before deployment, write down the tenant as the allocation unit, the entitlement source as the permission boundary, the denial status, the missing-rule behavior, and the fields allowed in telemetry. Fix the hash algorithm and revision the configuration. Then run the gate tests alongside the eval harness, using fixtures that include treatment, control, ineligible, and denied tenants.
During rollout, watch admitted request counts, denied reasons, error rates, latency, token usage, and cost units by revision and cohort. Compare quality results over the same slice. A rollback disables admission to the experimental handler; it should not rewrite historical cohort labels or erase the usage already attributed to that revision.
Finally, verify the privacy boundary with an actual event sample. The operational check is concrete: a reviewer should be able to explain experiment spend by tenant cohort and configuration revision, while being unable to reconstruct a student's prompt or lesson from the telemetry. That is enough observability to make the release decision, and little enough data to keep the gate understandable.
Further reading
- Martin Fowler, "Feature Toggles": https://martinfowler.com/articles/feature-toggles.html
- GDPR Article 5, "Principles relating to processing of personal data": https://gdpr-info.eu/art-5-gdpr/
Top comments (1)
tr.ee/dev-to