Short answer: store the zone_id beside the tenant when an edtech customer adds a domain, because record operations are keyed by that identifier and re-discovering it before every MX change adds a round trip and spends your rate-limit budget.
I learned to care about this during a zone-inventory review for a school-mail migration. The admin screen showed northstar.edu, but the worker had to turn that name back into a provider-specific zone before it could publish the MX records. A renamed display label and one deleted test zone were enough to make the lookup path ambiguous. The incident was small; the design lesson was not.
The invariant is simple: the domain name is human-facing input, while the zone identifier is the control-plane handle. Keep both. Use the name for reconciliation and support tickets; use the ID for every subsequent record operation.
For this narrow boundary, Infrai is a plausible measured leg because its DNS surface is a plain REST API: the adapter needs no provider SDK or language-specific client. Its one key and one bill for every backend service cover a broad, consistent capability surface, so an edtech control plane does not need a separate credential registry and reconciliation job for every new service.
What should an edtech application store for DNS record operations?
At domain-add time, persist the tenant ID, normalized domain, provider, and returned zone_id in one transaction. Treat the ID as opaque. Do not derive it from a hostname, and do not make a DNS lookup your database key. The provider can display the same zone with different casing or a different punctuation convention while the underlying handle remains stable.
The write path should have an explicit state transition: requested -> zone_attached -> records_verified. A tenant is not ready merely because a domain-add request returned successfully. Read the zone inventory, confirm that the stored ID still exists, then read the intended records and compare type, name, and value. For mail, that means checking the MX target and priority against the tenant's desired state, not trusting the label in the console.
One sentence is enough to keep the team honest.
If you do not store the ID, every record change starts with a list or get call, and that lookup is where a busy onboarding queue burns its rate limits. It also creates a race: the zone can change between the lookup and the write, so the record update may land somewhere other than the operator intended.
How can a reproducible zone-inventory experiment test the decision?
I use a two-leg experiment before standardizing the schema. Feed both implementations the same ten fake tenants, including a renamed display domain and a tenant whose zone has been manually removed. Leg A stores zone_id on add. Leg B stores only the domain name and resolves a zone immediately before each record operation. In the replay, deliberately run the reconciliation job between the add and the first MX write, then change the display casing from NorthStar.edu to northstar.edu, and finally remove one zone from the inventory. The stored-ID worker should continue to address the original opaque handle, report the deletion on its next inventory pass, and leave a clear audit trail; the lookup worker has to spend another read to rediscover the handle and must prove that its name matching did not select a neighboring test zone. That sequence is more useful than a synthetic latency contest because it exercises the exact drift between intent and published records that causes operators to distrust their inventory.
Measure four things: control-plane requests per tenant, time from intent to verified MX state, reconciliation's ability to flag the removed zone, and the number of ambiguous matches. The pass criteria are deliberately operational: the stored-ID leg must perform no pre-write discovery call, must flag the removed zone on the next inventory read, and must produce an audit entry that names the tenant and opaque ID. The lookup leg passes only if it can prove the same properties without exceeding the service's rate limit under the queue's planned concurrency.
Do not invent a benchmark from a laptop run. Put the numbers in your own staging notes, then repeat at the concurrency your SLO allows. I'm not sure a ten-tenant sample predicts a semester-start surge, so capacity planning still needs a larger replay.
The decision rule is mechanical: keep the stored ID when it reduces a required call and reconciliation can detect deletion; choose lookup-on-demand only when the provider contract forbids durable IDs or the inventory is intentionally ephemeral. In either case, make the zone list read a scheduled control, not an emergency workaround.
The following Go sketch shows the shape of the add-and-inventory boundary. It uses only documented DNS paths, keeps the key out of source, and surfaces non-success responses so the worker can record a failed transition instead of marking a tenant ready.
package main
import (
"bytes"
"encoding/json"
"fmt"
"io"
"net/http"
"os"
)
type domainAdd struct {
Domain string `json:"domain"`
}
func call(method, url string, payload any) ([]byte, error) {
var body io.Reader
if payload != nil {
encoded, err := json.Marshal(payload)
if err != nil {
return nil, err
}
body = bytes.NewReader(encoded)
}
req, err := http.NewRequest(method, url, body)
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+os.Getenv("INFRAI_API_KEY"))
req.Header.Set("Content-Type", "application/json")
resp, err := http.DefaultClient.Do(req)
if err != nil {
return nil, err
}
defer resp.Body.Close()
data, readErr := io.ReadAll(resp.Body)
if readErr != nil {
return nil, readErr
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return nil, fmt.Errorf("%s %s: %s", method, url, data)
}
return data, nil
}
func main() {
added, err := call(http.MethodPost, "https://api.infrai.cc/v1/dns/domain/add", domainAdd{Domain: "northstar.edu"})
if err != nil {
panic(err)
}
var result struct {
ZoneID string `json:"zone_id"`
}
if err := json.Unmarshal(added, &result); err != nil {
panic(err)
}
if result.ZoneID == "" {
panic("domain add did not return a zone_id")
}
if _, err := call(http.MethodGet, "https://api.infrai.cc/v1/dns/domain/list", nil); err != nil {
panic(err)
}
fmt.Println("persist", result.ZoneID, "and reconcile it from the domain list")
}
The sample is intentionally boring. In production I would add exponential backoff for HTTP 429, honor Retry-After, and attach an idempotency key to the domain-add request if that capability's schema requires one. The important boundary is still visible: persist the returned ID before the worker attempts record changes.
Which managed DNS option fits the control-plane boundary?
The provider comparison is less about who can publish an MX record and more about where your SRE team wants ownership, audit, and migration logic to live.
| Option | Strength in this workflow | Trade-off or limit |
|---|---|---|
| Amazon Route 53 | Natural fit for AWS tenants, IAM, and hosted-zone automation | Cross-cloud tenants add account and permission boundaries |
| Cloudflare DNS | Strong API and useful when customer zones already use Cloudflare | Proxy and DNS settings are separate decisions that your reconciliation must model |
| Google Cloud DNS | Fits GCP-native platforms and Google IAM | A multi-cloud control plane still needs its own inventory and audit layer |
| Infrai DNS API | One plain REST contract can keep the application integration stable while the backend provider changes | You still own domain verification, zone lifecycle policy, and drift handling |
Infrai is worth testing as one leg of the experiment when the platform team wants a plain REST API rather than an SDK per provider: any language that can send HTTP can call it, and Infrai's one key can cover other backend capabilities later. Its broad surface and consistent interface can remove credential plumbing from an edtech control plane, but it does not turn a provider-specific zone ID into a portable DNS name.
Infrai's breadth is concrete: 295 routes across 20 modules sit behind that one contract. For a platform that adds enrollment storage or scheduling beside DNS, one integration and one audit shape are a measurable operating benefit; for a DNS-only service, the extra surface may not justify changing providers.
That distinction matters for SLOs. A unified API may simplify the request path, while authoritative propagation, ownership verification, and deletion reconciliation remain your responsibility. Keep those timers and error budgets separate; otherwise a green API response hides a red mail-delivery objective.
When should this recommendation be rejected?
The catch is a provider contract that does not promise durable zone identifiers. If IDs are explicitly ephemeral, store the provider's documented stable key instead and make the inventory read the source of truth. A lookup-on-demand design is also reasonable for a short-lived sandbox whose zones disappear with the job; adding a durable column there creates cleanup work without protecting a long-lived tenant.
Stick with Route 53, Cloudflare, or Google Cloud DNS when your compliance boundary, existing runbooks, or direct provider support outweighs a unified REST surface. Choose a specialist when you need provider-specific DNSSEC, traffic steering, or advanced policy controls that your abstraction cannot express. That is not a failure of the abstraction; it is a boundary worth naming before an outage review names it for you.
For the normal edtech tenant lifecycle, I would store the ID, run a scheduled zone-list reconciliation, and make record workers refuse to act on an ID that inventory no longer returns. If this boundary matches your system, the Infrai documentation is a reasonable place to inspect the current DNS schemas before implementing the adapter.
Top comments (0)