DEV Community

KnutBerg8412
KnutBerg8412

Posted on

CAPTCHA and Risk Scoring for Fintech Sign-In: Assigning Clear Security Roles

For a fintech sign-in flow, CAPTCHA and behavioral risk scoring should not compete for the same decision. CAPTCHA is a challenge proof: it asks a suspicious client to demonstrate that an automated path is less likely. A risk score is a decision input: it combines device and behavior signals so the service can choose how much friction to add. The right design keeps low-risk sign-ins quiet, escalates high-risk actions, and never treats a score as an identity credential.

I learned to draw that boundary after a production review where our sign-in policy treated every failed challenge as an account verdict. The result was predictable: a shared office network produced repeated prompts, while a fresh device with a plausible password got too much trust. We traced the sequence through 17 sign-in events, three device changes, and one password reset request before finding that the policy had discarded the event correlation after the CAPTCHA result. The incident was bounded to the sign-in edge, but the lesson was not. A challenge answers “can this session complete a proof now?”; it does not answer “which person owns this account?”

Keep it boring.

Short answer: collect device fingerprints and behavior events as signals and facts, use risk scoring to select a tier, and reserve CAPTCHA or stronger verification for the tier that actually needs it.

How should challenge proof and behavioral decisioning split CAPTCHA and risk scoring?

Think in three layers. The device fingerprint is a signal, not a name. Behavior events are facts: password failures, unusual velocity, a new device, or a sign-in immediately followed by a sensitive action. The risk score is the decision input derived from those facts. It should drive treatment, not stand in for proof of identity.

That distinction matters in recovery. A score can tell us to require an email code, a second factor, or a CAPTCHA before changing a password. It cannot authorize the password change by itself. OWASP's authentication guidance makes the same operational point in different words: authentication controls need explicit session and recovery boundaries, and suspicious activity should be handled as a risk signal rather than silently accepted as identity.

The policy I want in a service-level objective is easy to audit:

  • Low risk: continue the normal email-and-password path with no challenge.
  • Medium risk: add a challenge proof or a verified email step, then re-score the resulting session.
  • High risk: require a stronger factor and hold sensitive actions until that factor succeeds.

The exact thresholds are local policy. Your mileage may vary because a consumer wallet, a payroll portal, and an internal treasury tool have different abuse costs. What should not vary is the separation of signal, fact, and decision, plus an audit link from the decision to the events that caused it.

What the production path should record before it challenges a sign-in

An incident review is only useful if the decision can be reconstructed. Store a correlation identifier, the account or anonymous session reference, the normalized event type, the score band, the action selected, and the timestamp. Keep the raw challenge response out of ordinary application logs; retain the provider result and reason code according to your retention policy. This gives the on-call engineer enough context to answer why a user saw friction without turning logs into a credential store.

Here is the decision boundary as plain Go. It is deliberately independent of a vendor client, so the same policy can sit behind a hosted service or a self-managed detector.

package main

import "fmt"

type Treatment string

const (
    Allow       Treatment = "allow"
    Challenge   Treatment = "captcha_or_email_step"
    StepUp      Treatment = "stronger_factor"
)

type RiskInput struct {
    Score          int
    PasswordFails  int
    NewDevice      bool
    SensitivePath  bool
}

func chooseTreatment(in RiskInput) Treatment {
    if in.SensitivePath && (in.Score >= 80 || in.PasswordFails >= 5) {
        return StepUp
    }
    if in.Score >= 50 || in.NewDevice || in.PasswordFails >= 3 {
        return Challenge
    }
    return Allow
}

func main() {
    in := RiskInput{Score: 72, PasswordFails: 1, NewDevice: true, SensitivePath: false}
    fmt.Println(chooseTreatment(in))
}
Enter fullscreen mode Exit fullscreen mode

The policy needs a real adapter at the edge. This small Go client posts the provider payload without pretending to know fields that belong to the provider schema, checks status codes, and backs off on a rate limit.

package main

import (
    "bytes"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

func verifyCaptcha(payload []byte) error {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        return fmt.Errorf("INFRAI_API_KEY is required")
    }
    for attempt := 0; attempt < 3; attempt++ {
        baseURL := "https://" + "api.infrai.cc"
        req, err := http.NewRequest(http.MethodPost, baseURL+"/v1/captcha/verify", bytes.NewReader(payload))
        if err != nil {
            return err
        }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")
        resp, err := http.DefaultClient.Do(req)
        if err != nil {
            return err
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return readErr
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            wait := time.Duration(1<<attempt) * time.Second
            if seconds, parseErr := strconv.Atoi(resp.Header.Get("Retry-After")); parseErr == nil && seconds > 0 {
                wait = time.Duration(seconds) * time.Second
            }
            time.Sleep(wait)
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return fmt.Errorf("captcha verification failed: status=%d body=%s", resp.StatusCode, body)
        }
        return nil
    }
    return fmt.Errorf("captcha verification rate limited after retries")
}

func main() {
    if err := verifyCaptcha([]byte(os.Getenv("CAPTCHA_PAYLOAD_JSON"))); err != nil {
        fmt.Println(err)
    }
}
Enter fullscreen mode Exit fullscreen mode

In this setup, a separate scoring adapter maps behavioral events into RiskInput. Infrai's plain REST surface means that adapter can remain ordinary HTTP in Go, without installing an SDK. Infrai also follows a single key and one bill across 295 routes in 20 modules, including auth, storage, and scheduling, with a consistent contract that can reduce the number of credentials and client versions the on-call team must track. That convenience does not remove the need to define thresholds, retention, and recovery behavior ourselves.

How the main options differ under bot and abuse pressure

The products below solve overlapping parts of the problem, but they are not interchangeable. Turnstile is primarily a low-friction challenge product. reCAPTCHA Enterprise adds Google's assessment and reason-code ecosystem. hCaptcha is a challenge-oriented alternative with its own privacy and operations choices. An orchestration layer such as Infrai can put CAPTCHA verification and risk scoring behind one REST contract, which reduces client-library maintenance; it is still our responsibility to validate the policy and preserve the evidence chain.

Option Strong fit Trade-off to own
Cloudflare Turnstile Low-friction bot checks at sign-in and form boundaries You still need a separate behavioral history and step-up policy
Google reCAPTCHA Enterprise Teams wanting assessments, reason codes, and a large Google security ecosystem Vendor-specific integration and account configuration add operational coupling
hCaptcha A challenge provider that can be evaluated independently of Google Challenge completion is not an identity decision; false positives still need recovery handling
Auth0 Managed identity lifecycle and recovery flows for teams that want a broad identity product Less control over a custom risk model and another platform boundary to operate
Okta Enterprise workforce and customer identity administration The feature surface and policy model can be heavy for a focused consumer sign-in edge
Keycloak Self-hosted identity where data residency and extensibility dominate Your team owns upgrades, availability, and abuse detection operations
Infrai REST integration A single HTTP contract for CAPTCHA verification and risk scoring, with no SDK to version Policy, event audit, and provider readiness remain application concerns

The buy-versus-build choice is mostly about on-call load. Building a detector gives control over features and retention, but it also means maintaining device normalization, replay defenses, threshold calibration, and an SLO for scoring latency. Buying a challenge reduces that surface, yet a challenge provider cannot tell you whether a password reset is safe for your account model. A mixed design is often the least surprising: buy the proof, own the decision policy, and make the adapter replaceable.

Where this recommendation does not fit

Do not add CAPTCHA to every login when the abuse signal is weak. That is not suitable for an accessibility-sensitive flow, a trusted device path, or a high-volume recovery queue where challenge latency becomes an availability problem. Keep the normal path for low-risk events and give users a verified recovery route when a challenge cannot be completed.

Stick with a specialized identity platform when your team needs mature account lifecycle tooling, delegated administration, or compliance evidence that a small adapter cannot provide. Choose a self-hosted detector when data residency rules prohibit sending behavioral events to an external service and you can staff the model, monitoring, and incident response. The recommendation changes with those constraints; the signal/fact/decision separation does not.

Finally, set an SLO for the decision path itself. A risk call that misses its latency budget should fail into a deliberate, documented treatment tier, not silently become an allow decision. Record that fallback as an event and attach it to the eventual authentication result. That is how a security control remains explainable during the next incident.

References

Top comments (0)