Short answer: put CAPTCHA before the first irreversible account-creation write, then apply risk scoring after signals arrive; measure each checkpoint separately so a support team can remove friction without removing the guardrail.
That is a placement rule, not a vendor choice. In a customer-support product, a new user may be trying to open a ticket while already frustrated. Making every visitor solve a challenge can lower abuse and legitimate completion at the same time. Letting every request create a user row, hash a password, and enqueue mail gives an automated campaign a convenient way to consume capacity. The boundary I use is the first side effect that is expensive or difficult to undo.
Why measure friction as a sequence instead of one signup metric?
Signup is a sequence of checkpoints: form load, challenge, provisional record, email verification, and first successful sign-in. A single “signup conversion” number hides where people leave. Track counters and latency for each transition, with dimensions that do not identify a person: coarse network class, client type, locale, and policy version. Keep raw identifiers out of analytics unless a documented abuse investigation needs them.
The useful signal is a pair of curves. Plot challenge completion against challenge presentation, then plot verified accounts against provisional records. If the first curve drops for screen-reader users but the second stays stable, the challenge experience needs work. If provisional records spike while verification stays flat, creation is too permissive or the mail path is being abused. Those are different interventions. In a support queue, this distinction changes the on-call move: a sudden rise in captcha_required during a browser release points to rendering or accessibility, while a rise in pending_verification from a narrow ASN range points to admission control. I want both counters on the same dashboard, with a link from the alert to a sampled trace, because paging an engineer to inspect a blended conversion metric at 3 a.m. is how a two-minute policy tweak becomes a two-hour incident review.
I initially treated a failed challenge as a generic authentication error. That made the support queue impossible to explain. A reason code such as captcha_missing, rate_limited, or risk_step_up is safer and more useful when it is stored server-side and shown only as a broad user message. Never turn it into an account-existence oracle.
Measure it.
Three words: observe before enforcing.
Start new scoring in shadow mode for at least one representative traffic window, including the 02:00 support lull and the Monday surge. Compare the proposed action with the current action, estimate the false-positive tail, and only then change who sees extra friction. Your mileage may vary across regions; a shared office network can look suspicious to a model even when every employee is legitimate.
How should CAPTCHA timing and risk scoring change a support signup?
Use CAPTCHA as an explicit gate before creation. The request can contain an email, password, and challenge response, but it should not yet create a fully active identity. Normalize only for comparison, enforce per-IP and per-email attempt budgets, and make retries idempotent. A challenge failure should consume a bounded attempt budget and return a generic response.
Once the challenge passes, create a provisional registration with an expiry and a one-time verification token. This record is deliberately boring: no ticket access, no agent invitations, and no password-reset privileges. Now the system has signals that did not exist at the edge, such as request velocity over several minutes, verification age, device continuity, and repeated targets in the same support workspace.
Risk scoring should select a next action, not pronounce a person good or bad. A low-risk registration can proceed to email verification. A high-risk one can require a second challenge, a longer cool-down, or manual review. Store the policy version, decision reason, and expiry so an on-call engineer can answer “why was I blocked?” without guessing which threshold changed.
The score call needs a deadline measured in milliseconds, a capped retry count, and a correlation ID. For an account-creation action, I prefer a bounded fail-closed response that leaves the provisional record pending and permits a later retry. For a password-reset email, a rate-limited queue may be more humane than blocking every request when a dependency is slow. The policy is action-specific.
Here is the shape of that decision in Go. The values are placeholders; calibrate them with your own false-positive and completion data.
package signup
type Decision struct {
CreatePending bool
RequireEmail bool
StepUp bool
Reason string
}
func choose(challengeOK bool, risk int, attempts int) Decision {
if !challengeOK {
return Decision{Reason: "captcha_required"}
}
if attempts > 5 {
return Decision{Reason: "rate_limited"}
}
if risk >= 80 {
return Decision{CreatePending: true, RequireEmail: true, StepUp: true, Reason: "risk_step_up"}
}
return Decision{CreatePending: true, RequireEmail: true, Reason: "pending_verification"}
}
Do not couple a password hash to a remote score with no timeout. Hashing consumes CPU; score retries can multiply that work during a bot burst. Put a small local budget in front of both operations and expose p95 and p99 latency, not only HTTP 200 counts. An SLO such as “99% of accepted signups reach provisional state within 2 seconds” makes the trade visible to the platform roadmap.
What failure modes appear after a support account is provisioned?
Support systems have attractive abuse targets: automated ticket floods, fake customer identities, invitation harvesting, and password-reset mail volume. A botnet that sends one request from each of 500 addresses can evade a per-IP limit. Aggregate by several privacy-reviewed keys, expire the data, and alert on combinations such as one device touching dozens of workspaces in 10 minutes.
Accessibility is part of the threat model. WCAG 2.2 requires more than a visual puzzle: keyboard access, readable status, and an alternative path matter when a legitimate customer uses assistive technology or a slow phone. Keep a non-CAPTCHA fallback behind stronger rate limits and email verification rather than silently exempting an entire network.
Replay is another quiet failure. A browser may retry after a dropped response, and an attacker may replay a valid challenge token. Use an idempotency key scoped to the registration attempt, reject an expired or already-consumed token, and ensure that sending verification mail has its own quota. The endpoint should converge on one provisional record, never on a pile of nearly identical accounts.
A buy-vs-build test for the friction checkpoints
| Checkpoint | Build in the application | Managed component boundary |
|---|---|---|
| Attempt budgets and idempotency | Domain keys, expiry, and replay rules belong here | Accept only atomic counters with documented TTLs |
| CAPTCHA presentation | Build the accessible form and fallback flow | A challenge service can provide the challenge; keep the user-facing policy local |
| Signal collection | Define privacy, retention, and workspace-level aggregation | A reputation feed may supply signals, not the final decision |
| Risk policy and audit | Keep thresholds, reason codes, and policy versions in your system | External scoring is an input with a strict timeout |
| Password handling | Use a reviewed, current hashing library | Hosted identity is acceptable only with tested export and recovery contracts |
The catch is ownership. A managed challenge reduces code and on-call surface, but it adds dependency latency, a privacy review, and rules your support staff may not be able to inspect. A self-hosted challenge offers control and repeatable tests, yet accessibility, key rotation, and abuse tuning become your 24-hour responsibility. Stick with the managed boundary when you cannot staff incident response; build more locally when you already operate the telemetry and review process.
Price is a secondary input. Count challenge calls, hash CPU, mail delivery, storage for pending records, and the engineer-hours required to investigate a false positive; a nominal per-call rate says little about the total SLO cost.
Verification and rollback runbook
Before rollout, test expired tokens, duplicate submissions, clock skew, challenge timeouts, and score responses arriving after their deadline. Run keyboard-only and screen-reader journeys. In load tests, model a bot burst separately from a genuine customer peak because the queues and cache keys behave differently.
Deploy policy changes behind a versioned flag. Observe first, then expose the new gate to a small cohort while comparing completion, verification, abuse reports, and p99 latency. Keep a kill switch for a signal or threshold; it must not disable password hashing, local rate limits, or email verification. Record every decision with its policy version so rollback is a controlled configuration change.
If the mail queue backs up, pause new verification sends and retain provisional records until their normal expiry. If challenge traffic rises sharply, protect the registration database with admission control and return 429 Too Many Requests after the local budget is exhausted. The user should see a retryable message, not a mysterious half-created identity.
Success is a bounded abuse rate with a healthy verified-signup rate, predictable tail latency, and support agents who can explain each friction checkpoint. Zero friction and zero abuse are incompatible goals; choose the point your error budget and customers can live with.
Top comments (0)