The page says that new users are landing in the wrong workspace just after a company hostname cutover. The visible failure is an account assignment, but the earlier signal lives in the domain registry: the DNS proof changed, the new mapping was published, and at least one lookup path still made a decision from stale state.
Short answer: auto-join only when a normalized email domain has exactly one active, verified workspace mapping; version that mapping, log the decision, and keep the previous version available until the observation window says rollback is no longer needed.
Do not guess.
This is a safety decision, not a string-matching convenience. A miss can go to manual selection or a pending state. An ambiguous match must never pick the first row. I have been paged by missed jobs and duplicate deliveries, so I use the same reflex here: make the operation idempotent, make ownership singular, and make replay boring.
What should the page have shown before the workspace routing alert?
The account-routing page fires late. By then, a user has supplied an email, the application has extracted its domain, and some registry or cache has returned a workspace. Work backward from that chain. The useful earlier warning is not merely "DNS changed." It is "the active routing generation is not consistently observable by the systems that can admit users."
Track four values together: the normalized email domain, workspace ID, verification state, and mapping generation. Then record the generation used for every admission decision. If the control plane says generation 18 is active while one decision path still reports generation 17, the operator has a precise rollback question. Without that generation, a wrong assignment and an old cache entry look like unrelated symptoms.
There are two distinct clocks. DNS proof visibility determines when a new ownership claim can be trusted. Application propagation determines when every lookup path uses the same accepted mapping. A fast hostname cutover does not collapse those clocks into one, and treating them as one is how a clean DNS change becomes a tenant-isolation incident.
Consider a hypothetical cutover from generation 17 to generation 18. The verification worker observes the expected proof and records generation 18 as a candidate, but generation 17 remains the only active mapping. Registry readers begin reporting which snapshot they have loaded. One reader still reports 17; two report 18 as available but inactive. At this point, a dashboard that only says "verified" is dangerously comforting because it hides both the active generation and reader convergence. The correct action is to wait, not to send a synthetic signup through the candidate and declare success. Once every admission reader acknowledges the candidate, activation changes one registry pointer. A new signup then records workspace ID and generation 18 together. If an alert follows, the responder can compare that stored generation with the registry history, reactivate generation 17, and stop new generation-18 admissions without pretending that already-created memberships vanished. This trace is intentionally specific about state and intentionally silent about elapsed minutes: no universal timing number is established here. What matters is that each transition has observable evidence, one owner, and a reversible action.
Wait for evidence.
The first instrumentation change I would make is a structured decision event with domain, workspace_id, mapping_generation, verification_state, and decision. Keep the email address itself out of routine telemetry when the domain is sufficient. Counters can then separate joined, no_match, ambiguous, and not_verified, while a gauge or inventory query shows which generation each lookup node has loaded.
How should an email lookup auto-join a user to the right workspace?
Start by parsing the address, not by splitting on the first @. Normalize the domain to the canonical form used when ownership was verified. Then query an index whose invariant is one active verified mapping per domain. The lookup returns a decision and the generation that produced it; account creation stores both in the same transaction as workspace membership, or performs an equivalent compare-and-set in systems without a relational transaction.
The interface below keeps DNS checking outside the request path. Verification is a control-plane job. Admission reads a published snapshot, which makes latency predictable and stops a temporary resolver result from changing the answer halfway through signup.
package routing
import (
"errors"
"net/mail"
"strings"
)
var (
ErrNoMapping = errors.New("no verified domain mapping")
ErrAmbiguous = errors.New("multiple active domain mappings")
)
type Mapping struct {
Domain string
WorkspaceID string
Generation uint64
Verified bool
Active bool
}
type Registry interface {
FindActiveByDomain(domain string) ([]Mapping, error)
}
type Decision struct {
WorkspaceID string
Generation uint64
}
func LookupWorkspace(email string, registry Registry) (Decision, error) {
address, err := mail.ParseAddress(email)
if err != nil {
return Decision{}, err
}
at := strings.LastIndexByte(address.Address, '@')
if at < 1 || at == len(address.Address)-1 {
return Decision{}, ErrNoMapping
}
domain := strings.ToLower(strings.TrimSuffix(address.Address[at+1:], "."))
matches, err := registry.FindActiveByDomain(domain)
if err != nil {
return Decision{}, err
}
verified := make([]Mapping, 0, len(matches))
for _, mapping := range matches {
if mapping.Active && mapping.Verified && mapping.Domain == domain {
verified = append(verified, mapping)
}
}
switch len(verified) {
case 0:
return Decision{}, ErrNoMapping
case 1:
return Decision{
WorkspaceID: verified[0].WorkspaceID,
Generation: verified[0].Generation,
}, nil
default:
return Decision{}, ErrAmbiguous
}
}
The example deliberately does not infer a parent company from subdomain text, follow MX records, or treat email authentication as workspace ownership. DMARC defines an organizational-domain concept for mail policy, but a mail-policy boundary is not automatically your authorization boundary. The mapping registry needs its own explicit ownership proof and tenant policy.
Auto-join should also be idempotent. Replaying the same request for the same user, workspace, and mapping generation must return the existing membership instead of creating a second side effect. If a user already belongs to another workspace, route the case to an explicit transfer or review flow. Silent reassignment is too dangerous.
Replay must be dull.
The cutover is a versioned state transition
A safe cutover has an old generation, a candidate generation, and a recorded activation point. Verification of the candidate does not have to activate it immediately. That separation lets an operator prove control, warm the read path, and inspect consistency before new users are admitted under the candidate.
| Phase | Admission behavior | Rollback posture |
|---|---|---|
| Verify | Continue using the old active generation | Remove the candidate without changing users |
| Observe | Resolve the candidate in checks, but do not auto-join with it | Keep the old generation authoritative |
| Activate | Use the candidate only after lookup nodes report it loaded | Atomically reactivate the old generation |
| Retire | Stop accepting new decisions from the old generation | Preserve its audit record and explicit recovery procedure |
The key order is publish, observe, activate. During activation, reject ambiguity before accepting speed. A lookup that sees both generations as active should produce ambiguous, not choose the numerically larger generation and hope that every writer agrees.
Rollback follows the same model in reverse. Reactivate a known verified generation, publish it, and observe that admission paths have loaded it. Do not delete the candidate as the first response; deletion destroys the evidence needed to understand which decisions were made during the window. Memberships already created under the candidate require a separate review policy because rolling back routing state does not automatically prove those memberships are wrong.
This is where propagation delay competes with cutover speed. Shortening every observation window makes the change look fast, but it transfers uncertainty to user admission. Extending the window reduces that uncertainty while delaying automatic enrollment. There is no universal duration in the available evidence; your mileage may vary with resolver behavior, cache topology, and how quickly every admission node reports its generation. The go/no-go rule should therefore use observed convergence, not a duration copied from another system.
Failure handling belongs in the product contract
Three outcomes are routine and should not page anyone: no verified mapping, an invalid email address, and a user who needs manual workspace selection. Ambiguity is different. It breaks the registry invariant and should block auto-join immediately, create a high-signal operator event, and preserve enough context to identify both mappings.
The catch is that automatic company-domain enrollment is not suitable for organizations where contractors, schools, holding companies, or shared mail domains span independent security boundaries. In those cases, stick with invitations, administrator approval, or an identity-provider assertion whose tenant binding is explicit. A verified domain proves control of the verification mechanism; it does not prove that every mailbox holder should receive the same access.
That boundary matters.
Keep the user-facing fallback plain. "We could not select a workspace automatically" is accurate for a miss. It doesn't expose registry internals, and it avoids turning a cautious refusal into an apparent authentication failure. For operators, the runbook should ask which generation made the decision, whether exactly one mapping was active, whether its proof remained valid under policy, and whether all readers had acknowledged the generation.
Tests should cover mixed-case domains, a trailing dot in stored canonical data, malformed addresses, zero matches, one verified match, an unverified row, duplicate active rows, and replay of the membership write. Add a cutover test that holds one reader on the old generation while another loads the candidate. The expected result before activation is boring: both continue to admit against the old active mapping. After activation, the lagging reader should be removed from admission service until it acknowledges the new generation rather than making a stale decision.
Alert on decisions, then pay the false-positive bill
The primary service-level signal is the outcome of admission decisions, partitioned by mapping generation. A sudden rise in no_match after activation can reveal an incomplete publish. Any ambiguous result indicates a violated invariant. A count of readers on non-active generations tells the operator about propagation before users report the symptom.
One page is enough if it is actionable: page on ambiguity or on confirmed wrong-workspace decisions, and attach the two competing mappings plus the decision generation. Use a ticket or a lower-urgency alert for ordinary misses unless their rate changes materially around a cutover. I'm not sure a fixed miss-rate threshold transfers between products; baseline traffic, signup campaigns, and expected personal-email usage would resolve that uncertainty.
Thresholds have a cost. If a single no_match pages the on-call, normal typos and personal addresses will train responders to distrust the alert. If the threshold waits for a large batch, the system may assign many users before anyone acts. Tie urgency to security impact and state-transition timing: ambiguity is immediate, generation skew blocks activation, and a miss-rate change is evaluated against a local baseline. The final check in the runbook should be whether the page identifies an action the responder can take. If it cannot, change the signal before widening the pager rotation.
References
- RFC 7489, Domain-based Message Authentication, Reporting, and Conformance (DMARC): https://datatracker.ietf.org/doc/html/rfc7489
Top comments (0)