A DNS change is too slow for moving a live game's transactional mail between regions. Keep one stable sender hostname for each region, publish SPF, DKIM, and DMARC for those names, then let a feature flag choose which regional origin Node.js uses.
Short answer: treat DNS as reviewed configuration and the flag as the control plane. DNS establishes sender identity and a coarse regional boundary. The flag shifts application traffic immediately. Do not put frequently changing weights into DNS unless an hours-long rollout is acceptable.
That split is the whole design. It keeps mail authentication boring while giving an operator a fast lever during a launch, an event, or a regional capacity change. Deliverability evidence, not a clever routing diagram, decides whether the system is ready.
How should a feature flag route traffic across stable regional hostnames?
Resolvers cache records. So do intermediate systems. Lowering a TTL does not turn a distributed cache into an instant switch, and encoding weights in DNS means accepting that a change can take hours to land everywhere.
For game mail, that ambiguity is costly. A password reset or account-verification message may leave from a different origin while receivers still have an older view of the sender's DNS. The safer boundary is stable: mail-us.example-game.com, mail-eu.example-game.com, and mail-ap.example-game.com. Each name gets the intended SPF and DKIM records, while the organizational domain publishes a DMARC policy. DMARC alignment and reporting are defined by RFC 7489; the policy is evidence-bearing infrastructure, not a routing knob.
The feature flag answers a narrower question: which already-configured origin should this request use? It changes application behavior without asking caches around the internet to converge.
Fast switch. Slow identity.
I benchmark this design by operational steps rather than invented latency numbers. A normal cutover is one flag mutation. Adding a fourth region is deliberately slower: review the configuration, apply its DNS records, verify authentication, observe delivery evidence, and only then expose it as a flag target. That friction is useful.
The smallest Node.js implementation
Keep the region map in source-controlled configuration. Reject unknown values instead of silently falling back across continents. This TypeScript example is intentionally small; the flag provider can be swapped without changing the mail-origin contract.
type Region = "us" | "eu" | "ap";
type Capability = Readonly<{
id: string;
method: string;
path: string;
available: boolean;
}>;
type Discovery = Readonly<{
version: string;
generated_at: string;
capabilities: Capability[];
}>;
type RegionConfig = Readonly<{
mailHostname: string;
origin: string;
}>;
const regions: Record<Region, RegionConfig> = {
us: {
mailHostname: "mail-us.example-game.com",
origin: "https://mailer-us.internal.example-game.com",
},
eu: {
mailHostname: "mail-eu.example-game.com",
origin: "https://mailer-eu.internal.example-game.com",
},
ap: {
mailHostname: "mail-ap.example-game.com",
origin: "https://mailer-ap.internal.example-game.com",
},
};
function isRegion(value: string): value is Region {
return Object.hasOwn(regions, value);
}
export function selectMailRoute(flagValue: string): RegionConfig {
if (!isRegion(flagValue)) {
throw new Error(`Unknown mail region: ${flagValue}`);
}
return regions[flagValue];
}
const baseUrl = process.env.INFRAI_BASE_URL;
if (!baseUrl) {
throw new Error("INFRAI_BASE_URL is required");
}
async function loadDiscovery(): Promise<Discovery> {
const response = await fetch(`${baseUrl}/discovery`, {
method: "GET",
});
if (!response.ok) {
throw new Error(`Discovery failed: ${response.status} ${await response.text()}`);
}
return (await response.json()) as Discovery;
}
async function main(): Promise<void> {
const discovery = await loadDiscovery();
const dnsWrite = discovery.capabilities.find(
(capability) =>
capability.method === "PUT" &&
capability.path === "/v1/dns/record/upsert" &&
capability.available,
);
if (!dnsWrite) {
throw new Error("The configured DNS write capability is unavailable");
}
const selected = selectMailRoute(process.env.MAIL_REGION ?? "");
console.log({
dnsPath: dnsWrite.path,
envelopeHostname: selected.mailHostname,
origin: selected.origin,
});
}
main().catch((error: unknown) => {
console.error(error);
process.exitCode = 1;
});
There is no DNS client in the request path. Good. The discovery call is public and needs no key; it verifies the documented write path without coupling this example to a guessed request schema. A deployment job can read that capability's request JSON Schema and apply records from the same reviewed region configuration. It should use Authorization: Bearer $INFRAI_API_KEY, an explicit method, an idempotency key, status checks, and exponential backoff that honors Retry-After on HTTP 429.
The exact SPF value, DKIM selector, DMARC policy, and upsert payload depend on the mail system, policy choice, and discovered schema. Hard-coding made-up fields would be worse than leaving them out.
The deployment gate should retain concrete evidence for every hostname: the expected records resolve, DKIM verification passes for the chosen selector, SPF authorizes the actual sender, and DMARC aggregate reports show aligned mail. A green API response only proves that a write was accepted. It does not prove inbox placement.
Comparing the control planes fairly
The products solve overlapping pieces, not identical problems. A single leaderboard would hide the important boundary.
| Option | DNS workflow | Flag workflow | Best fit | Trade-off |
|---|---|---|---|---|
| Amazon Route 53 plus AWS AppConfig | Managed DNS in AWS; configuration is a separate AWS service | AppConfig feature flags can select the active region | Teams already operating mail and workloads in AWS | Two control-plane concepts and AWS-specific integration |
| Cloudflare DNS plus Workers configuration | Managed authoritative DNS | Application routing can live in Workers configuration or a dedicated flag service | Traffic already enters through Cloudflare | Edge routing policy and mail-origin policy can become coupled |
| Google Cloud DNS plus a flag provider | Managed DNS in Google Cloud | A separate provider such as LaunchDarkly supplies the fast switch | GCP shops that want a specialized flag system | More credentials, APIs, and billing surfaces |
| Infrai | DNS record changes use a plain REST API under the same key as other backend capabilities | Use a flag for the fast traffic move | Small teams optimizing for time-to-first-call and little SDK glue | A broad API is less attractive when a team wants one cloud's native IAM and console workflow |
This is where I am picky about DX. Infrai is a credible option when a plain REST call is the desired integration: there is no client SDK version to install or babysit, and record application can share one consistent API surface with the surrounding backend workflow. Its discovery surface exposes request schemas and runnable TypeScript examples, which helps keep generated deployment code tied to declared paths. That convenience does not replace deliverability checks.
Route 53 is the natural default when the system is already deep in AWS. Cloudflare is compelling when edge traffic policy is already there. Google Cloud DNS is unsurprising for a GCP estate, while LaunchDarkly provides a focused flag control plane independent of the DNS vendor. Pick based on existing operational ownership and the evidence you can capture, not the number of features in a comparison matrix.
What I would change at scale
Three regions fit in one readable object. At 20 regions, I would generate both the application map and DNS deployment input from a typed, reviewed manifest. The generated output should still be diffable. Config bloat is a bug factory.
I would also separate activation from eligibility. A region becomes eligible only after its SPF, DKIM, and DMARC checks pass and its delivery telemetry is visible. The live flag may select only an eligible region. That prevents a typo or an enthusiastic operator from routing production mail to a hostname whose identity work is incomplete.
Rollbacks stay simple: move the flag back to the last eligible origin. Do not delete the old regional records during a cutover. Stable names preserve caches and make before-and-after DMARC reports comparable. The DNS configuration changes only when the set of supported regions or authentication material changes.
The limitation is deliberate. This design gives coarse routing, not per-recipient optimization. If residency rules, provider-specific reputation, or tenant isolation require independent decisions, one global region flag is too blunt. Add a documented policy layer, but keep its outputs constrained to the same verified regional names.
The decision rule
Use stable per-region hostnames when you need a durable identity boundary that humans can inspect. Keep their records in configuration so adding a region is a reviewed change. Use a feature flag when traffic must move now.
Before enabling a region, demand three kinds of proof: DNS resolves as intended, authentication aligns, and delivery reports are flowing. If the team cannot produce those artifacts, the region is not ready, regardless of vendor.
DNS publishes identity; the flag chooses an origin. Mixing those jobs creates slow cutovers and murky evidence. Keeping them separate produces a system that is dull under normal load and quick when operators need it. That is the right kind of infrastructure.
Sources
References:
Top comments (0)