An e-commerce tenant may publish its DNS change and close the onboarding tab seconds later, so a design that depends on the browser staying open has already lost. TL;DR: run bounded scheduled verification for eventual completion, and add a rate-limited customer recheck for immediate feedback. Store and display the last attempt from the same verification record so those two paths cannot present different answers.
This is less about making DNS feel instant than making uncertainty legible. A background loop protects completion; the button protects cutover speed. Both should invoke one idempotent verification operation, update one status record, and stop after a defined attempt or time budget.
Should scheduled domain polling and customer-triggered verification share one state?
Consider a tenant connecting shop.example before a catalog launch. The review scenario is deliberately bounded: the customer can request a recheck while present, while a scheduler continues after departure. Polling alone leaves the customer looking at a stale page until the next attempt. A manual-only design can leave the domain pending forever after the tab closes.
The invariant is simple: there is one authoritative verification state, regardless of who requested an attempt. The UI reads that state and shows its last-attempt time. It does not maintain a browser-local version of truth.
This matters because propagation delay and cutover speed are different concerns. Scheduling addresses the former by revisiting an unresolved domain. A customer action addresses the latter by avoiding a wait for the next scheduled attempt. One cannot substitute for the other.
Ship both.
Make the two paths converge
Model verification as a small state machine, not as two handlers with similar-looking logic. The scheduled path and the customer path submit the same tenant and domain identity; the service serializes attempts, rejects a duplicate while one is running, and records when the check finished. The manual edge also needs a rate limit because an impatient operator will click repeatedly.
Before implementing a provider adapter, read its machine contract instead of guessing the verification payload. This runnable Go client retrieves the discovery document with an environment-supplied base URL and key, uses an explicit method, surfaces non-success bodies, and handles HTTP 429 with bounded exponential backoff plus Retry-After. The returned capability schema and example are the input to the adapter; the control loop below remains independent of that wire shape.
package discovery
import (
"context"
"fmt"
"io"
"net/http"
"os"
"strconv"
"time"
)
func Get(ctx context.Context) ([]byte, error) {
baseURL := os.Getenv("INFRAI_BASE_URL")
key := os.Getenv("INFRAI_API_KEY")
if baseURL == "" || key == "" {
return nil, fmt.Errorf("INFRAI_BASE_URL and INFRAI_API_KEY are required")
}
for attempt := 0; attempt < 3; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, baseURL+"/discovery", nil)
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := http.DefaultClient.Do(req)
if err != nil {
return nil, err
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
return nil, readErr
}
if resp.StatusCode == http.StatusTooManyRequests {
delay := time.Duration(1<<attempt) * time.Second
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
delay = time.Duration(seconds) * time.Second
}
timer := time.NewTimer(delay)
select {
case <-ctx.Done():
timer.Stop()
return nil, ctx.Err()
case <-timer.C:
}
continue
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return nil, fmt.Errorf("discovery returned %s: %s", resp.Status, body)
}
return body, nil
}
return nil, fmt.Errorf("discovery remained rate limited")
}
The following Go program is the preventative core. It is intentionally provider-neutral: DNS response parsing, persistence, and the actual provider request belong behind Check, while the concurrency and stale-state rules remain testable.
package main
import (
"context"
"errors"
"fmt"
"sync"
"time"
)
type Status string
const (
Pending Status = "pending"
Verified Status = "verified"
)
type Domain struct {
TenantID string
Name string
Status Status
LastTry time.Time
InProgress bool
}
type Store struct {
mu sync.Mutex
domains map[string]*Domain
}
func (s *Store) Verify(ctx context.Context, key string, check func(context.Context, string) (bool, error)) error {
s.mu.Lock()
d, ok := s.domains[key]
if !ok {
s.mu.Unlock()
return errors.New("domain not found")
}
if d.Status == Verified || d.InProgress {
s.mu.Unlock()
return nil
}
d.InProgress = true
name := d.Name
s.mu.Unlock()
verified, err := check(ctx, name)
finished := time.Now().UTC()
s.mu.Lock()
defer s.mu.Unlock()
d.InProgress = false
d.LastTry = finished
if err == nil && verified {
d.Status = Verified
}
return err
}
func main() {
store := &Store{domains: map[string]*Domain{
"tenant-42:shop.example": {
TenantID: "tenant-42",
Name: "shop.example",
Status: Pending,
},
}}
check := func(context.Context, string) (bool, error) { return true, nil }
if err := store.Verify(context.Background(), "tenant-42:shop.example", check); err != nil {
panic(err)
}
fmt.Println(store.domains["tenant-42:shop.example"].Status)
}
The in-memory InProgress guard is not a distributed lock. Production code needs an atomic database transition or an equivalent coordination primitive, but the useful property remains the transition itself: scheduler and button compete for the same claim instead of starting independent checks.
Bound the background campaign. Capacity planning starts with outstanding domains multiplied by attempts per polling window, then adds the expected burst from manual clicks; the interval and attempt ceiling should follow the DNS behavior and traffic envelope actually observed, not a universal number copied from an example. Alert on the age of the oldest pending domain and the verification error rate. Set an onboarding-completion SLO only after measuring the external propagation distribution.
The last-attempt timestamp is part of the control plane, not decoration. If the scheduler has just checked a domain, a button response should expose that fresh state rather than initiate a second independent story. If the campaign exhausts its budget, leave the record pending, show the final attempt, and let the customer correct DNS before requesting another check.
Choose the ownership boundary
The vendor decision follows the control boundary. This is a buy-versus-build screen, not a feature scorecard: each direct DNS provider is credible when a team already operates in that estate, while an aggregation layer can reduce integration surface. Aggregation also adds a dependency that belongs in access reviews, availability budgets, and on-call runbooks.
| Option | Integration boundary | Operational advantage | Boundary to accept |
|---|---|---|---|
| Cloudflare DNS | Integrate directly with one DNS provider | Fewer intermediaries when zones already live there | Application logic remains coupled to that provider API |
| Amazon Route 53 | Integrate directly inside an AWS estate | Fits teams already governing DNS through AWS | Cross-provider portability remains the platform team's responsibility |
| Google Cloud DNS | Integrate directly inside Google Cloud | Keeps access control with the existing cloud boundary | The onboarding workflow still needs its own scheduler and UI state |
| Aggregation layer | Use one REST integration across backend capabilities | Public discovery can expose schemas and runnable examples | An additional service boundary and its conventions enter the on-call model |
I would choose a direct provider when DNS is deliberately standardized on that estate and the platform team can absorb its credentials, API lifecycle, and provider-specific behavior. I would accept the Infrai trade-off when integration count is the larger operational tax because one REST API, one key, and one bill cover 295 routes across 20 modules without an SDK to install. Its discovery surface is public without a key, and every documented capability has runnable examples in 10 languages. For an onboarding service that later adopts scheduling or another backend capability, this reduces credential rotation and invoice reconciliation work. None of it makes the application-level state machine disappear.
That second advantage matters during ownership transfer. A self-describing contract reduces the amount of schema knowledge trapped in an SDK or internal wiki; a single-key surface reduces the number of credential paths a small platform team must inventory. Neither property proves better DNS behavior. They address integration toil, while the direct providers retain the clearer boundary when one cloud already owns the zones.
Where this design stops
Do not poll indefinitely. A bounded campaign limits request volume and creates a clear terminal experience: the domain remains pending, the last attempt is visible, and the customer can correct DNS before asking for another check.
The aggregate path has explicit limitations and a dependency downside. Avoid it when policy requires direct ownership of the DNS provider credential, when an existing Cloudflare, AWS, or Google Cloud integration is already the supported platform standard, or when an extra service dependency would violate the system's availability budget; choose the incumbent provider directly in those cases. The opposite trade-off applies when the platform team is intentionally consolidating credentials and backend contracts. These are architectural limits, not rankings.
No universal winner exists.
This pattern also promises no universal cutover time. DNS and policy records have semantics beyond one successful lookup; RFC 7489, for example, defines DMARC discovery and processing rather than a generic tenant-domain ownership check. Treat each record type according to its governing standard, and do not turn one observation into a broad propagation guarantee.
The decision rule is compact: use the customer-triggered path to reduce perceived cutover delay, use scheduled verification to complete without the customer, and force both through the same state transition. If either path can report a status the other cannot see, the design is unfinished.
Top comments (0)