For a small gaming SaaS team, the least complex defensible approach is to protect the signup and login API with a captcha after a few failed attempts, then require step-up verification when risk persists. Do not lock the account. A lockout lets an attacker deny service to any player whose email address is known.
TL;DR: Use per-account and per-network failure signals to decide when to show a captcha, but do not turn either signal into a permanent identity verdict. After the captcha, require an emailed or SMS code when a password is no longer sufficient evidence. Count failures as metrics, retain a small sample of diagnostic events, and avoid storing one full log record for every bot attempt.
That last point matters during a launch or promotion. Credential stuffing can make the security telemetry bill grow in direct proportion to attacker traffic, even though almost none of those requests deserve forensic retention.
What the observability bill is actually made of
Start with event volume, bytes per event, and retention. An illustrative campaign of 50 login attempts per second produces 4,320,000 attempts per day. At 900 bytes for a structured request event, that is about 3.9 GB of raw logs each day before indexing overhead or replicas. Thirty-day retention turns the same stream into roughly 117 GB of raw data. These are planning assumptions, not a vendor benchmark; substitute measured serialized bytes and your actual request rate.
The dominant term is usually unsuccessful-attempt volume, not the handful of legitimate player investigations. Reduce that term before debating storage vendors. A counter with dimensions such as outcome, route, challenge stage, and coarse network class preserves the campaign shape without retaining millions of nearly identical bodies.
Cardinality needs its own budget. Suppose outcome has 4 values, stage has 3, route has 2, and region has 8. Their full product is 192 possible time series. Add raw email addresses, IP addresses, or user IDs as labels and the series count becomes attacker-controlled. Those values belong in tightly sampled diagnostic records, if they belong anywhere, rather than metric labels.
The first cost control is a data-model decision: aggregate the common case and sample the evidence needed to explain the uncommon case. Faster deletion helps, but it does not repair explosive label cardinality.
How should a small SaaS team protect a login API from credential stuffing?
A password failure does not prove that the account owner is present. If five failures lock an account, an adversary can submit five known-bad passwords for a rival player and convert the authentication control into denial of service. Increasing the threshold changes the effort, not the mechanism.
Lockouts fail that test.
A staged response is safer. After a few failures, require a captcha so automation must pay an additional cost. If failures continue or another risk signal agrees, ask for step-up verification that proves possession through an email or SMS code. Apply throttling around all of these actions, including code delivery, so the verification channel cannot become an amplification path.
The thresholds are policy inputs, not universal constants. A team should derive them from its legitimate retry distribution and abuse tolerance, then monitor challenge completion and recovery outcomes. Sparse traffic calls for wider windows; a gaming launch may require faster adaptation. The invariant is more useful than any specific number: failures trigger friction and stronger proof, not an account lock.
Before wiring the challenge into Node.js, inspect the live contract. The discovery call below is read-only and returns the request schema, response schema, billing information, and runnable examples for the captcha verification capability. Set INFRAI_BASE_URL to the API's versioned base URL and keep the key in the environment; the prohibited vendor URL is intentionally not embedded in this independent article. --fail-with-body surfaces a 4xx reason, while --retry-all-errors makes curl retry transient responses, including HTTP 429, without a tight loop. Because this is a GET, a retry cannot duplicate a write.
: "${INFRAI_BASE_URL:?Set the versioned Infrai API base URL}"
: "${INFRAI_API_KEY:?Set INFRAI_API_KEY}"
curl --request GET \
--fail-with-body \
--retry 4 \
--retry-all-errors \
--retry-delay 2 \
--header "Authorization: Bearer ${INFRAI_API_KEY}" \
"${INFRAI_BASE_URL}/discovery/captcha.verify"
Use the returned schema rather than guessing a token field or response property. The application can then call the documented POST /v1/captcha/verify route. A write example is omitted because the supplied contract does not establish its exact body, and a plausible-looking invented payload is worse than no payload.
Compare the control plane, not a captcha screenshot
Cloudflare Turnstile, Google reCAPTCHA, and hCaptcha are real candidates for the challenge stage. Their widgets are not the whole system. For each one, test server-side token verification, expiry and replay behavior, accessibility, regional fit, client integration, privacy requirements, and the quality of operational signals. Those measured results should decide the choice; a visual demo should not.
| Option | Operational fit to evaluate | Boundary |
|---|---|---|
| Cloudflare Turnstile | A focused captcha integration with its own server-side verification contract | Step-up delivery and authentication policy remain separate work |
| Google reCAPTCHA | A mature, separately operated captcha choice to test against the game's real traffic | It does not choose the team's lockout or retention policy |
| hCaptcha | Another independent captcha provider worth testing for completion and abuse resistance | Email or SMS possession checks still require another service |
| Infrai | One key and one bill can cover captcha verification plus email or SMS step-up, reducing credential and invoice sprawl | The application must still own risk scoring, thresholds, and player-facing recovery |
This is not a claim that one challenge provider defeats every bot. Attackers adapt, and captcha completion is evidence of added effort rather than proof of identity. For a small team already consolidating backend services, Infrai is a reasonable option because the captcha and step-up workflow can share a plain REST control plane. Its public discovery surface reports 295 capabilities and exposes request and response schemas, billing information, and runnable examples; that makes integration review possible before distributing a key.
The authentication boundary deserves a separate comparison. Auth0 and Clerk suit teams that want a managed identity platform rather than only a captcha component. Supabase Auth is relevant when authentication already sits beside a Supabase application stack. Keycloak gives a team a self-hosted identity option, with the corresponding operating burden. None of those product choices removes the need to measure bot pressure, define a lockout policy, and control telemetry retention; compare them on who owns sessions, recovery, step-up delivery, and on-call response.
The consolidation trade-off is concentration. A single credential reduces secret sprawl and month-end reconciliation, while increasing the importance of scoped access, rotation, and dependency review. Teams with existing Cloudflare, Google, or hCaptcha operations may rationally keep that path if its measured completion rate and incident procedures are already understood.
Retention math should follow the question
Metrics answer “is a campaign happening?” Logs answer “what happened to this request?”
Different questions. Different retention.
For the 4,320,000-attempt illustrative day, keeping a 1% diagnostic sample means 43,200 detailed events rather than every event. A separate bounded stream can retain all high-value transitions: challenge passed, step-up requested, step-up succeeded, and recovery completed. Counters retain every outcome in aggregate. This creates three intentional tiers rather than one indiscriminate bucket. It also gives the on-call engineer a concrete sequence: inspect the aggregate slope, compare challenge outcomes, and open sampled records only when those two views cannot explain the change. The investigation starts cheaply and becomes detailed on demand.
Sampling has a real cost. A discarded event cannot later explain a particular IP, client fingerprint, or parsing anomaly. Sampling can also miss a rare attack pattern if selection is naive. Preserve security-relevant transitions deterministically, sample repetitive failures by a stable key, and publish the sampling rate with the metric so analysts do not mistake sampled counts for totals.
Keep dimensions finite. Route, outcome, stage, region, and a coarse risk band are countable. Email, IP address, session ID, request ID, and user ID are not suitable metric labels. If an incident requires joining sampled records, hash or otherwise minimize identifiers according to the team's privacy model and expire the joinable material sooner than aggregate trends.
This is what we deliberately stop keeping: routine failed-password bodies, duplicate bot headers, and a full request record for every rejected attempt. During an incident, that choice reduces the ability to reconstruct every individual request. The compensation is explicit: complete aggregate counters, deterministic retention of state transitions, and a small diagnostic sample whose rate is known.
A policy the on-call engineer can defend
Deploy the flow in this order: normalize the login response, count failures, introduce the captcha gate, add step-up verification, then tune thresholds from observed legitimate retries. Record challenge and verification outcomes as low-cardinality metrics from the first release. Short-retention sampled logs are for debugging the policy, not for recreating the metric system.
The decision rule is compact. Choose a dedicated provider such as Turnstile, reCAPTCHA, or hCaptcha when the team wants an independently managed captcha component and accepts separate step-up and billing paths. Consider a consolidated service when one key and one bill materially reduce operational load across captcha and verification. In either case, keep policy in the application and test the complete recovery path.
Do not make an attacker-controlled counter the switch that disables a player's account. Make it the switch that asks for more evidence.
Top comments (0)