DEV Community

Sam for Adios.dev

Posted on Originally published at adios.dev

How to Build an Anycast CDN with BIRD, NGINX, and GeoDNS

To build an Anycast CDN, announce the same IP prefix from multiple locations and run an HTTPS cache at each one.

In this guide, we'll use:

  • BIRD for BGP
  • NGINX for HTTPS and caching
  • IPv4 /24 routing
  • RPKI and ROAs
  • Health-based route withdrawal
  • GeoDNS for regional Anycast pools

We'll start with two servers sharing one global CDN address, then expand the design into regional pools.


Choose global Anycast or regional Anycast with GeoDNS

Global Anycast uses one service IP across all your edges.

DNS returns that IP, and BGP selects the receiving edge according to network policy and available paths.

A cache hit is served there. A cache miss goes to your origin and can populate that edge's cache.

Regional Anycast with GeoDNS adds another selection step.

Each regional pool has a different service IP shared by the edges in that pool. GeoDNS returns a regional IP, and BGP then selects an edge advertising it.

This lets you choose a regional pool before internet routing chooses the server.

Routing choice Global Anycast Regional Anycast + GeoDNS
DNS answer The same CDN IP for everyone. A different CDN IP for each selected regional pool.
BGP announcements All edges announce the same prefix. Edges in each pool announce that pool's prefix.
One edge fails Withdraw its route; other announcing edges remain available. Withdraw its route; another edge in the same pool can receive new connections.
An entire region fails Another advertising region may receive the traffic after convergence. DNS must select a healthy fallback pool. Cached DNS answers can still point to the failed pool.
Choose it when You want one stable IP and one global pool. You want explicit regional pool selection, separate capacity, or different regional origins.

The request flow looks like this:

GLOBAL ANYCAST

cdn.example.com -> one IP -> BGP -> Paris or New York cache
                                              |
                                          cache miss
                                              v
                                            origin


REGIONAL ANYCAST + GEODNS

cdn.example.com -> GeoDNS -> Europe IP -> BGP -> Paris / Frankfurt
                         -> US IP     -> BGP -> New York / Chicago
                         -> default   -> chosen fallback pool
Enter fullscreen mode Exit fullscreen mode

A regional pool is a deployment choice.

Its routes can still be reachable globally. Restricting where BGP announcements propagate is a separate upstream policy.

Neither GeoDNS nor BGP guarantees the nearest server or enforces data residency by itself.

What you need for the first build

Start with the global setup.

You'll need:

  • Two fresh Debian 12 servers with systemd
  • BIRD 2
  • NGINX
  • Customer BGP support from both providers
  • One authorized IPv4 /24
  • An agreed origin ASN
  • An HTTPS origin on a separate IP
  • A domain whose DNS you control

The examples below assume direct BGP peers.

Your upstream must provide the real peering values.

Budget for:

  • Address space
  • ASN arrangements
  • Edge servers
  • Outbound traffic
  • Origin traffic
  • DNS health checks
  • Monitoring

Get current quotes and confirm BGP eligibility before ordering infrastructure.

All IP addresses and ASNs shown in this article are documentation examples and must be replaced with your own.

Deployment and failure timings must be measured on your own network.

References


Get a /24 and permission to announce it

For your own IPv4 announcements on the public internet, /24 is the practical minimum block size.

A /24 contains 256 addresses.

Higher prefix lengths represent smaller networks:

  • /25 = 128 addresses
  • /26 = 64 addresses

Those smaller blocks are commonly filtered on the public internet.

Before buying or leasing address space, confirm that both hosting locations support customer BGP and will accept your prefix.

There are several ways to obtain a /24.

Route What to do Check before committing
Lease a block with your own ASN Lease a /24 and have the holder authorize your ASN to originate it. Arrange the ASN separately if needed. Confirm ROA and IRR updates, permission to announce from both PoPs, and renewal and exit terms.
Request a provider-assigned block Ask a provider such as Vultr for address space it can route for your BGP setup. Confirm whether a full /24 is available, which ASN originates it, and whether you can announce it elsewhere.
Apply directly to a registry Request an allocation from your regional internet registry. RIPE NCC maintains a /24 waiting list for eligible LIRs that have never received an IPv4 allocation. Membership or application fees do not guarantee allocation or delivery time.
Buy an existing block Purchase a block from an existing holder directly or through a broker and complete the registry transfer process. Verify seller authority, transfer eligibility, fees, routing history, abuse history, and upstream acceptance.

Provider-managed Anycast is another entry point.

A provider can expose individual Anycast service IPs from inside its own aggregate.

AWS Global Accelerator, for example, provides managed Anycast IPs and supports eligible BYOIP ranges while AWS handles the announcements.

That is different from operating customer BGP yourself.

Make the block routable

Agree the origin ASN with your upstreams.

If you need your own ASN, apply through your regional registry or through an eligible sponsoring provider.

The ASN identifies the network originating the prefix. It is separate from the IP allocation, lease, or transfer.

The resource holder should authorize the ASN using RPKI.

Create a ROA for the /24 with:

Maximum length: /24
Enter fullscreen mode Exit fullscreen mode

Add the corresponding IRR route object if required by your upstream, and provide a Letter of Authorization if requested.

Remember:

A ROA authorizes a route. It does not announce the route.

Send something like this to both hosting providers:

Our prefix:       [your /24]
Our origin ASN:   [your ASN]
Locations:        Paris and New York, announced simultaneously

Please confirm:
- Prefix accepted; required ROA, IRR record, and authorization
- Peer IP, peer ASN, local source IP, and any BGP password
- Direct or multihop peering
- Replies sourced from our /24 are allowed
- Global export policy and available regional communities
Enter fullscreen mode Exit fullscreen mode

Do not configure production BGP until both providers confirm that they will accept the prefix and provide their peer configuration.

References


Prepare the first edge server

We'll put the first edge in Paris and the second in New York.

Each location is a Point of Presence, or PoP.

Both locations receive the same CDN IP.

Each server also keeps its own normal unicast address for:

  • SSH
  • Management
  • Origin requests
  • Health monitoring

Replace all addresses, ASNs, and example.com names below.

These are documentation values, not addresses you can announce.

Our example address plan is:

cdn.example.com → 203.0.113.80
                       |
                  BGP chooses a PoP
                   /             \
              Paris cache     New York cache
                   \             /
                    cache misses
                        |
               origin.example.com
                  192.0.2.10:443

Prefix:               203.0.113.0/24
CDN IP on BOTH PoPs:  203.0.113.80/32
Origin ASN:           64496
Upstream ASN:         64497
Paris node / peer:    198.51.100.10 / 198.51.100.1
New York / peer:      198.51.100.20 / 198.51.100.17
Enter fullscreen mode Exit fullscreen mode

Install packages and bind the CDN address

Run this on the Paris server:

sudo apt-get update
sudo apt-get install -y bird2 nginx curl ca-certificates iproute2 python3
sudo install -d -o www-data -g www-data /var/cache/nginx/cdn
Enter fullscreen mode Exit fullscreen mode

The CDN service IP will be bound as a /32.

A covering blackhole route will discard packets sent to unused addresses in the /24.

Allow:

  • HTTPS traffic to the CDN IP
  • BGP TCP port 179 from your provider's peer

Keep management traffic on the server's unicast IP.

This reverse-proxy configuration does not require Linux packet forwarding.

Keep the address across reboots

Create:

/etc/systemd/system/cdn-address.service
Enter fullscreen mode Exit fullscreen mode

with:

[Unit]
Description=CDN service address
Before=bird.service nginx.service

[Service]
Type=oneshot
RemainAfterExit=yes
ExecStart=-/usr/sbin/ip link add anycast0 type dummy
ExecStart=/usr/sbin/ip address replace 203.0.113.80/32 dev anycast0
ExecStart=/usr/sbin/ip link set anycast0 up
ExecStart=/usr/sbin/ip route replace blackhole 203.0.113.0/24

[Install]
WantedBy=multi-user.target
Enter fullscreen mode Exit fullscreen mode

Order services after the address

Make BIRD and NGINX depend on this unit.

for service in bird nginx; do
  sudo mkdir -p /etc/systemd/system/$service.service.d

  sudo tee /etc/systemd/system/$service.service.d/cdn-address.conf >/dev/null <<'EOF'
[Unit]
Requires=cdn-address.service
After=cdn-address.service
EOF
done

sudo systemctl daemon-reload
sudo systemctl enable --now cdn-address

ip address show dev anycast0
ip route get 192.0.2.10
Enter fullscreen mode Exit fullscreen mode

The route to the origin should use the normal unicast network, not anycast0.

References


Configure BIRD to announce the /24

Create:

/etc/bird/bird.conf
Enter fullscreen mode Exit fullscreen mode

on the Paris server.

router id 198.51.100.10;

protocol device {
  scan time 10;
}

protocol static cdn_prefix {
  disabled yes;

  ipv4;

  route 203.0.113.0/24 blackhole;
}

filter export_cdn {
  if net = 203.0.113.0/24 then accept;
  reject;
}

protocol bgp transit {
  local as 64496;
  source address 198.51.100.10;
  neighbor 198.51.100.1 as 64497;

  graceful restart off;

  ipv4 {
    import none;
    export filter export_cdn;
  };
}
Enter fullscreen mode Exit fullscreen mode

The important part is the export filter.

Only the CDN prefix can be exported.

filter export_cdn {
  if net = 203.0.113.0/24 then accept;
  reject;
}
Enter fullscreen mode Exit fullscreen mode

The cdn_prefix static protocol starts disabled.

That means the server initially advertises nothing.

We will only advertise the route after HTTPS has been tested locally.

Check the BGP session

Add any provider-required authentication or multihop options.

Then validate the configuration:

sudo bird -p -c /etc/bird/bird.conf
sudo systemctl enable --now bird
sudo birdc configure

sudo birdc show protocols all transit
sudo birdc show route export transit
Enter fullscreen mode Exit fullscreen mode

The BGP session should reach:

Established
Enter fullscreen mode Exit fullscreen mode

but no CDN route should be exported yet.

If the BGP session does not establish, check:

  • Peer address
  • Peer ASN
  • Local ASN
  • Firewall rules
  • Authentication
  • Multihop requirements

Reference


Issue HTTPS certificates

Use DNS-01 certificate validation so you can issue the CDN certificate before advertising the CDN IP.

Here is a Certbot example for a zone hosted on Cloudflare DNS.

Install:

sudo apt-get install -y certbot python3-certbot-dns-cloudflare
sudo install -d -m 700 /root/.secrets
sudo touch /root/.secrets/cloudflare.ini
sudo chmod 600 /root/.secrets/cloudflare.ini
sudoedit /root/.secrets/cloudflare.ini
Enter fullscreen mode Exit fullscreen mode

Inside:

dns_cloudflare_api_token = YOUR_TOKEN
Enter fullscreen mode Exit fullscreen mode

Restrict the Cloudflare token to the required DNS zone with:

Zone:DNS:Edit
Enter fullscreen mode Exit fullscreen mode

Then request the certificate:

sudo certbot certonly --dns-cloudflare \
  --dns-cloudflare-credentials /root/.secrets/cloudflare.ini \
  --dns-cloudflare-propagation-seconds 60 \
  --cert-name cdn.example.com \
  -d cdn.example.com \
  --email ops@example.com \
  --agree-tos \
  --non-interactive
Enter fullscreen mode Exit fullscreen mode

Issue the certificate independently on each edge.

The servers do not have to share the same private key.

As the edge fleet grows, however, you may want to centralize issuance and securely distribute certificates to reduce DNS credential exposure.

Reload NGINX after renewal

After the NGINX configuration later in this guide passes nginx -t, install a Certbot deploy hook:

sudo install -d /etc/letsencrypt/renewal-hooks/deploy

sudo tee /etc/letsencrypt/renewal-hooks/deploy/reload-nginx >/dev/null <<'EOF'
#!/bin/sh
set -eu

/usr/sbin/nginx -t
/usr/bin/systemctl reload nginx
EOF

sudo chmod 755 /etc/letsencrypt/renewal-hooks/deploy/reload-nginx

sudo systemctl enable --now certbot.timer
sudo certbot renew --dry-run --run-deploy-hooks

systemctl list-timers certbot.timer
Enter fullscreen mode Exit fullscreen mode

References


Configure HTTPS and the edge cache

We'll use an HTTPS origin at:

192.0.2.10:443
Enter fullscreen mode Exit fullscreen mode

and an origin hostname of:

origin.example.com
Enter fullscreen mode Exit fullscreen mode

The origin should serve:

/assets/logo.v1.svg
Enter fullscreen mode Exit fullscreen mode

with headers similar to:

Cache-Control: public, max-age=60, s-maxage=300
ETag: "cdn-demo-v1"
Enter fullscreen mode Exit fullscreen mode

It should also serve:

/cdn-probe.txt
Enter fullscreen mode Exit fullscreen mode

with:

Cache-Control: no-store
Enter fullscreen mode Exit fullscreen mode

Create:

/etc/nginx/conf.d/cdn.conf
Enter fullscreen mode Exit fullscreen mode

with:

proxy_cache_path /var/cache/nginx/cdn
  levels=1:2
  keys_zone=cdn:50m
  max_size=10g
  inactive=60m
  use_temp_path=off;

map "$http_authorization$http_cookie" $skip_private {
  default 1;
  ""      0;
}

upstream cdn_origin {
  server 192.0.2.10:443;
  keepalive 32;
}

server {
  listen 203.0.113.80:443 ssl;
  server_name cdn.example.com;

  ssl_certificate     /etc/letsencrypt/live/cdn.example.com/fullchain.pem;
  ssl_certificate_key /etc/letsencrypt/live/cdn.example.com/privkey.pem;
  ssl_protocols TLSv1.2 TLSv1.3;

  if ($host != cdn.example.com) {
    return 421;
  }

  proxy_http_version 1.1;
  proxy_set_header Connection "";
  proxy_set_header Host origin.example.com;
  proxy_set_header X-Forwarded-Proto https;
  proxy_set_header X-Forwarded-For $remote_addr;

  proxy_ssl_server_name on;
  proxy_ssl_name origin.example.com;
  proxy_ssl_verify on;
  proxy_ssl_trusted_certificate /etc/ssl/certs/ca-certificates.crt;
  proxy_ssl_verify_depth 3;

  proxy_connect_timeout 3s;
  proxy_read_timeout 15s;

  proxy_hide_header X-Edge-Id;
  proxy_hide_header X-Cache;

  add_header X-Edge-Id "paris-1" always;
  add_header X-Cache $upstream_cache_status always;

  location = /__edge/health {
    default_type text/plain;
    return 200 "edge-ok\n";
  }

  location ^~ /assets/ {
    proxy_pass https://cdn_origin;

    proxy_cache cdn;
    proxy_cache_key "$scheme|$host|$request_uri";
    proxy_cache_methods GET HEAD;

    proxy_cache_bypass
      $skip_private
      $http_cache_control
      $http_pragma;

    proxy_no_cache
      $skip_private
      $http_cache_control
      $http_pragma
      $upstream_http_set_cookie;

    proxy_cache_lock on;
    proxy_cache_revalidate on;
  }

  location / {
    proxy_pass https://cdn_origin;
  }
}
Enter fullscreen mode Exit fullscreen mode

This configuration only caches /assets/.

Everything else goes directly to the origin.

Give the origin predictable test responses

If the origin also uses NGINX, add these locations inside its HTTPS server:

location = /assets/logo.v1.svg {
  default_type image/svg+xml;

  add_header Cache-Control "public, max-age=60, s-maxage=300";
  add_header ETag '"cdn-demo-v1"';

  return 200 '<svg xmlns="http://www.w3.org/2000/svg" width="80" height="80"><rect width="80" height="80" fill="blue"/></svg>';
}

location = /cdn-probe.txt {
  default_type text/plain;
  add_header Cache-Control "no-store";

  return 200 "origin-ok\n";
}
Enter fullscreen mode Exit fullscreen mode

What this cache will store

The test asset remains fresh in the shared cache for 300 seconds because of:

s-maxage=300
Enter fullscreen mode Exit fullscreen mode

The NGINX setting:

inactive=60m
Enter fullscreen mode Exit fullscreen mode

controls removal of unused objects, not HTTP freshness.

Our cache key includes:

  • Scheme
  • Hostname
  • Full request URI

Requests with either:

  • Cookies
  • Authorization headers

bypass the cache.

NGINX also respects relevant:

  • private
  • no-store
  • Set-Cookie
  • Vary

response behavior.

Each PoP has its own disk cache.

Use versioned asset URLs when releasing new versions.

This basic configuration does not provide:

  • A global purge API
  • Forced stale serving during origin failure
  • A shared global cache

References


Enable the first PoP

Before announcing anything to the internet, test HTTPS locally.

sudo nginx -t
sudo systemctl enable --now nginx
sudo systemctl reload nginx
Enter fullscreen mode Exit fullscreen mode

Test the edge health endpoint:

curl --fail --show-error --max-time 3 \
  --resolve cdn.example.com:443:203.0.113.80 \
  https://cdn.example.com/__edge/health
Enter fullscreen mode Exit fullscreen mode

Expected:

edge-ok
Enter fullscreen mode Exit fullscreen mode

Then test origin connectivity:

curl --fail --show-error --max-time 5 \
  --resolve cdn.example.com:443:203.0.113.80 \
  https://cdn.example.com/cdn-probe.txt
Enter fullscreen mode Exit fullscreen mode

Expected:

origin-ok
Enter fullscreen mode Exit fullscreen mode

Both checks must succeed with a valid certificate before announcing the route.

Announce the route and set DNS

Enable the BIRD static protocol:

sudo birdc enable cdn_prefix
sudo birdc show route export transit
Enter fullscreen mode Exit fullscreen mode

Ask your upstream to confirm the /24 is accepted and propagated.

Then create:

cdn.example.com. 300 IN A 203.0.113.80
Enter fullscreen mode Exit fullscreen mode

If your authoritative DNS provider also offers a CDN proxy, keep this record in DNS-only mode.

Test the hostname from an external network before proceeding.


Add the second PoP

Repeat the setup in New York.

Keep these values the same:

  • /24
  • CDN service IP
  • Origin ASN
  • CDN hostname
  • Origin
  • Cache configuration

Change the server-specific BGP and edge values.

Setting Paris New York
BIRD router ID 198.51.100.10 198.51.100.20
BGP source address 198.51.100.10 198.51.100.20
BGP neighbor 198.51.100.1 198.51.100.17
NGINX X-Edge-Id paris-1 new-york-1

Install a valid certificate and verify local HTTPS before enabling cdn_prefix.

Confirm both locations serve requests

Leave DNS unchanged.

Probe the public hostname from multiple networks and inspect:

X-Edge-Id
Enter fullscreen mode Exit fullscreen mode

Depending on internet routing policy, clients should reach different PoPs.

To test a specific edge, run curl --resolve directly from that server.

Using the Anycast IP from your laptop does not let you force the request to Paris or New York.


Withdraw unhealthy edges automatically

A live BGP session does not mean HTTPS is healthy.

An edge might still be advertising while:

  • NGINX is dead
  • TLS is broken
  • The service address disappeared
  • The application is stuck

The edge therefore needs health-aware route withdrawal.

The example controller used by this guide:

  1. Probes the local CDN IP with the correct TLS hostname
  2. Withdraws cdn_prefix after three consecutive failures
  3. Requires 30 seconds of continuous health before recovery
  4. Uses a 60-second withdrawal hold-down
  5. Reads BIRD state back after every action

Download the example files:

curl --fail --show-error --remote-name \
  https://adios.dev/examples/anycast-cdn/edge-health.py

curl --fail --show-error --remote-name \
  https://adios.dev/examples/anycast-cdn/edge-health.service
Enter fullscreen mode Exit fullscreen mode

Review both files before installing them.

Then:

sudo install -m 755 edge-health.py /usr/local/sbin/edge-health.py
sudo install -m 644 edge-health.service /etc/systemd/system/edge-health.service

sudo tee /etc/default/edge-health >/dev/null <<'EOF'
CDN_HOST=cdn.example.com
CDN_IP=203.0.113.80
EOF

sudoedit /etc/default/edge-health

sudo systemctl daemon-reload
sudo systemctl enable --now edge-health

sudo journalctl -u edge-health -n 30 --no-pager
Enter fullscreen mode Exit fullscreen mode

The systemd unit attempts route withdrawal when the controller exits and restarts it if it crashes.

It also uses a watchdog to detect a stalled controller loop.

This example checks:

  • Local TLS
  • NGINX responsiveness

You should separately monitor:

  • Cache disk health
  • Origin reachability
  • Public routing
  • Upstream connectivity
  • Route propagation

A shared origin outage should not necessarily withdraw every edge that can still serve cached content.

Test the controller on your own network.

The policy and command handling may have automated tests, but that does not mean your BGP environment is automatically production-qualified.

Test process failure, hung controllers, BIRD socket errors, upstream graceful restart behavior, and reboots.

References


Test cache hits and failover

On each PoP, request the same test asset twice.

for attempt in 1 2; do
  curl --silent --show-error --fail --max-time 10 \
    --resolve cdn.example.com:443:203.0.113.80 \
    -D - \
    -o /dev/null \
    https://cdn.example.com/assets/logo.v1.svg
done
Enter fullscreen mode Exit fullscreen mode

A fresh cache key should show:

MISS
Enter fullscreen mode Exit fullscreen mode

followed by:

HIT
Enter fullscreen mode Exit fullscreen mode

Check the origin access log.

The second request should not reach the origin.

An initial HIT simply means the object was already cached.

Now test a private request:

curl --silent --show-error --fail --max-time 10 \
  --resolve cdn.example.com:443:203.0.113.80 \
  -H 'Cookie: session=cdn-test' \
  -D - \
  -o /dev/null \
  https://cdn.example.com/assets/logo.v1.svg
Enter fullscreen mode Exit fullscreen mode

This should report:

X-Cache: BYPASS
Enter fullscreen mode Exit fullscreen mode

Withdraw one PoP

From an external network currently reaching Paris, repeatedly create new HTTPS connections.

On Paris, stop the automatic controller first:

sudo systemctl stop edge-health
Enter fullscreen mode Exit fullscreen mode

Then withdraw the route:

sudo birdc disable cdn_prefix
sudo birdc show route export transit
Enter fullscreen mode Exit fullscreen mode

External requests should eventually begin reaching New York.

Record:

  • Failed requests
  • Total failover time
  • Edge ID
  • Connection behavior

Existing TCP connections may need to reconnect.

Once local HTTPS is healthy again:

sudo systemctl start edge-health
Enter fullscreen mode Exit fullscreen mode

The controller should require sustained health before advertising the route again.

Test a service failure and a reboot

With the route advertised and controller running, stop NGINX on Paris:

sudo systemctl stop nginx
Enter fullscreen mode Exit fullscreen mode

Wait for the failure threshold.

Inspect:

sudo journalctl -u edge-health -n 30 --no-pager
sudo birdc show route export transit
Enter fullscreen mode Exit fullscreen mode

The health controller should withdraw the prefix.

External requests should begin reaching New York.

Restart NGINX:

sudo systemctl start nginx
Enter fullscreen mode Exit fullscreen mode

Verify that the recovery delay prevents immediate re-advertisement.

Repeat the same experiment on the other edge.

Then reboot one edge while the other continues serving traffic.

Verify the following recover in the expected order:

  1. Dummy interface
  2. Anycast service address
  3. BIRD
  4. BGP session
  5. NGINX
  6. Certificate availability
  7. Health controller
  8. Route advertisement

Measure both:

  • Warm-cache behavior
  • Cold-cache behavior

Never assume a fixed BGP convergence time.

Measure it on your own providers and client networks.

Reference


Add regional Anycast pools with GeoDNS

Once the global Anycast setup works, you can create regional pools.

Build at least two edges in each pool.

For example:

Europe

EU_CDN_IP
   |
   +-- Paris
   |
   +-- Frankfurt
Enter fullscreen mode Exit fullscreen mode

North America

US_CDN_IP
   |
   +-- New York
   |
   +-- Chicago
Enter fullscreen mode Exit fullscreen mode

Every edge in a regional pool advertises the same pool prefix.

Each independently routed IPv4 pool needs an independently routable prefix.

In practice, that normally means another /24.

Two different IP addresses from one shared /24 do not give you two independently controllable BGP routes.

For each pool, change:

  • Dummy interface address
  • Covering route
  • BIRD static prefix
  • BIRD export filter
  • NGINX listen address
  • Health controller CDN_IP

Keep management and origin IP addresses outside those service prefixes.

All pools should serve the same hostname with valid certificates and matching cache behavior.

You can use separate origins per region if required.

An example GeoDNS policy:

Query location DNS answer for cdn.example.com Edges announcing that IP
Europe EU_CDN_IP Paris and Frankfurt using EU_PREFIX/24
North America US_CDN_IP New York and Chicago using US_PREFIX/24
Default / fallback GLOBAL_CDN_IP Global fallback pool using GLOBAL_PREFIX/24

Create the GeoDNS records

With a GeoDNS platform such as Route 53, create multiple geolocation A records with the same hostname:

cdn.example.com
Enter fullscreen mode Exit fullscreen mode

but different record identifiers.

For example:

Europe        -> EU_CDN_IP
North America -> US_CDN_IP
Default       -> GLOBAL_CDN_IP
Enter fullscreen mode Exit fullscreen mode

Start with a TTL around:

60 seconds
Enter fullscreen mode Exit fullscreen mode

and measure actual resolver behavior.

If you move authoritative DNS from Cloudflare to Route 53, remember that the Cloudflare Certbot DNS plugin no longer controls the authoritative records.

Either:

  • Use the Route 53 DNS plugin, or
  • Deliberately delegate the ACME challenge zone

Associate each GeoDNS record with the health of its regional pool.

With Route 53, an unhealthy geographic record can fall back to:

  1. A broader geographic record
  2. The default record

Keep the fallback pool healthy and large enough to absorb additional traffic.

Check pool health, then test both failure paths

You need two kinds of checks.

First, monitor individual edges through their unicast addresses.

Second, probe the shared regional Anycast service IP.

A probe against the Anycast IP may continue succeeding after one server fails because BGP sends it to another healthy edge.

Therefore:

Remove a regional pool from GeoDNS when the pool itself cannot serve traffic, not whenever a single edge fails.

Test two failure scenarios.

Scenario 1: one edge fails

Withdraw Paris.

Frankfurt should continue serving the European service IP.

Scenario 2: an entire regional pool fails

Make the European pool unavailable in a controlled test.

Fresh DNS responses should select the fallback pool.

Also test clients that still have the old DNS answer cached.

DNS changes cannot:

  • Move an existing TCP connection
  • Instantly replace every cached DNS answer

GeoDNS also does not know the user's exact location.

It estimates based on:

  • Recursive resolver location
  • EDNS Client Subnet hints where available

Test from multiple real networks and public resolvers.

Also remember:

A default DNS record is not a guaranteed fail-safe.

Route 53 can return unhealthy records when all eligible records fail health checks.

Testing from target regions might look like:

dig +short cdn.example.com A

curl --silent --show-error --fail --max-time 10 \
  -D - \
  -o /dev/null \
  https://cdn.example.com/assets/logo.v1.svg
Enter fullscreen mode Exit fullscreen mode

Record:

DNS answer
X-Edge-Id
X-Cache
HTTP errors
Latency
Enter fullscreen mode Exit fullscreen mode

Repeat during:

  • One-edge withdrawal
  • Entire pool failure
  • Recovery

Other routing setups

GeoDNS isn't the only option.

You can also have globally reachable nodes and ask an upstream to limit route propagation using documented BGP communities.

This is a separate routing policy from GeoDNS.

NO_EXPORT, for example, refers to AS boundaries. It does not mean "keep this route inside Europe."

If a hosting provider cannot provide customer BGP, another option is to use a transit provider that announces your prefix and delivers traffic over tunnels.

If you do that, test:

  • Return routing
  • Source filtering
  • MTU
  • Path MTU discovery

And place an actual cache at the receiving location.

Sending every regional tunnel back to one distant cache leaves the central cache and its network path in every request.

References


Keep the CDN working

Building the first two PoPs is only the beginning.

Monitor each location through its unicast management path as well as through the public Anycast IP.

Otherwise, a failed server can disappear from your monitoring because the Anycast probe simply moves to another healthy location.

Here are some of the most important failure modes.

Problem What you will notice What to do
GeoDNS points clients at an unavailable pool Fresh lookups or cached answers continue reaching a failed region. Check pool health, fallback policy, DNS caching, and fallback capacity. Test whole-region failure separately from one-edge withdrawal.
Lease, ROA, or routing-record changes Some networks stop reaching the prefix even though BGP sessions remain established. Track renewal dates and RPKI validity. Recheck authorization before changing ASN or upstream.
Traffic shifts after a provider change Clients reach a distant PoP or overload a smaller location. Measure X-Edge-Id and latency from several access networks. Review provider routing policy before adjusting communities or prepending.
One PoP fails or its cache restarts The surviving edge or origin receives a large traffic spike. Test failover capacity with a cold cache and reserve disk/bandwidth headroom. Add an origin shield if duplicate fills become expensive.
Old or private content is cached Users receive stale releases or another user's response. Use versioned assets and test Cookie, Authorization, private, no-store, and Set-Cookie behavior.
A deployment or certificate differs between PoPs Only some networks see TLS failures, errors, or old behavior. Roll out one PoP first. Track certificate expiry and configuration versions per node.
The health controller fails or flaps A bad edge remains advertised or traffic repeatedly switches locations. Supervise the controller, use recovery hold-downs, test watchdog behavior, and coordinate graceful restart with upstreams.
A tunnel or return path breaks Small requests work but larger transfers stall, or responses never arrive. Check MTU, PMTU discovery, return routing, and source-address filtering.
An attack saturates the link Servers are technically healthy but unreachable before HTTP limits can help. Arrange upstream DDoS mitigation and understand how to activate it. Restrict origin access and monitor egress.
Several services share the same /24 Withdrawing the prefix for one service also moves healthy services. Define health policy for every service on the prefix or use separate prefixes where independent withdrawal is required.

Your first working CDN

You're ready to add real traffic when:

  • Both PoPs serve valid HTTPS
  • Each location proves MISS followed by HIT
  • Private responses stay uncached
  • Withdrawing either route moves new connections to the surviving location
  • Origin connectivity is monitored
  • The health controller has been failure-tested
  • Route recovery has been tested
  • Reboots have been tested
  • External probes confirm real internet routing behavior

Start with public static assets.

Measure the effect on origin traffic before expanding what you cache.

From there you can add:

  • More PoPs
  • More regional pools
  • Origin shielding
  • Purge APIs
  • Centralized certificates
  • Deployment automation
  • Configuration versioning
  • DDoS protection
  • Edge rate limiting
  • Capacity-aware routing
  • Regional origins

The difficult part of Anycast isn't getting two servers to announce the same address.

The difficult part is keeping routing, application health, TLS, cache behavior, failure recovery, and operational state synchronized as the network grows.

References


About Adios.dev

This article comes from the engineering work behind Adios.dev, where we're building a developer cloud for running applications, APIs, databases, and AI workloads.

You can explore the network here:

If you're building your own Anycast or regional CDN and want help with BGP, routing, caching, monitoring, upgrades, or ongoing operation, you can also reach the Adios engineering team.


Originally published on Adios.dev.

Top comments (0)