DEV Community

Cover image for Threat-model the model endpoint your app exposes to the internet
Conor Bronsdon
Conor Bronsdon

Posted on Originally published at chainofthought.show

Threat-model the model endpoint your app exposes to the internet

You ship a classifier, a moderation layer, or an agent that calls a model on every user request. Treat that endpoint as a machine-learning surface with a different abuse model than a CRUD API. Kris Lovejoy argues attackers can now find individual vulnerabilities faster than ever, and chain low-severity issues into high-impact paths. She also expects a fully autonomous AI-driven attack against an enterprise network within roughly a year and a half of her June 2026 episode. Whether or not that timeline lands, the direction is clear: more automation on the attack side means your production model call needs a threat model alongside its latency budget.

I’d start with five attacks: four that NIST groups under adversarial machine learning, plus model denial of service from OWASP’s list, using plain definitions from the Chain of Thought glossary and one control your app team can own for each.

Did someone poison the model before it reached you?

A backdoor attack compromises the model during training so it looks fine on ordinary inputs and switches behavior when it sees a trigger: a phrase, a pattern, or a small patch on an image. Clean accuracy tests will not catch it, because the failure mode is conditional. Outsourced training, third-party weights, and opaque fine-tunes are the supply-chain version of this risk.

Your control: treat every non-base model artifact (adapter, LoRA bundle, vendor fine-tune) like untrusted code. Pin provenance, document who trained on what, and run trigger-style red-team prompts before you promote a build. You will not “spot the backdoor” with a happy-path eval alone.

Can users slip past your filter at inference time?

An evasion attack does not change the weights. The attacker perturbs an input the model should block (spam, fraud, policy-violating content) until the model misclassifies it while a human still recognizes the intent. NIST ties this to adversarial examples; for LLMs, jailbreaking sits in the same family. Attackers often need only query access, including transfer from another model.

Your control: normalize and constrain inputs before they hit the model (length limits, canonicalization, block known perturbation patterns where you can), and keep a human-readable audit trail when the model clears borderline content. Adversarial training is a training-team problem; your job is to avoid giving the model a single brittle gate with no fallback.

Does your API leak training data through the response shape?

A model inversion attack reconstructs information about training data by querying the model and analyzing outputs. Fredrikson and colleagues showed that confidence scores returned alongside predictions can be enough to recover sensitive attributes or recognizable faces. The fix pattern is to return less.

Your control: default the public API to labels or coarse scores instead of full probability vectors or rich logprobs. If product needs scores internally, serve them on a separate, authenticated path with stricter rate limits and logging. Rounding or bucketing probabilities is boring and effective.

Can an attacker learn who was in the training set?

Membership inference exploits the gap between how a model behaves on records it saw during training versus novel records. In some domains, membership alone is sensitive (for example, participation in a study). Shokri and colleagues demonstrated the attack against commercial ML services; Carlini and colleagues later showed language models can regurgitate verbatim training snippets under query access.

Your control: shrink the inference contract. Return top-1 or top-k classes instead of full distributions where possible, add per-identity query budgets, and flag sessions that systematically probe many similar records. Training-time regularization and differential privacy are upstream; your endpoint still should not hand attackers a high-precision loss oracle.

Can one client drain your GPU budget or your bank account?

Model denial of service covers classic availability hits and “denial of wallet”: unbounded inference that burns compute or pay-per-token spend. OWASP’s LLM Top 10 lists unbounded consumption, including floods of variable-length inputs and inputs that repeatedly exceed the context window. Agent loops that fan out into many model and tool calls raise the same cost risk. Sponge-style inputs can spike latency and energy use on language models.

Your control: enforce budgets outside the model: max input tokens, max output tokens, rate limits per key, queue depth caps, timeouts, and for agents a hard cap on steps and tool calls per task. Monitor spend and latency anomalies the way you monitor 5xx rates.

Why frame it this way now?

Lovejoy’s worry extends past external actors. Agentic workflows are outcome-oriented; guardrails written for linear human processes miss parallel tool use. Insiders with utility access can chain weaknesses you classified as low risk because you optimized for crown jewels.

what we know now is that, you know, attackers are able to find individual vulnerabilities faster than we ever have before. And so that means that we have to get better at detection and better at remediation.

Kris Lovejoy, Global Head of Strategy at Kyndryl, on Chain of Thought ep 62

Now we've gotten to a point where the stuff that we didn't worry about is a big problem.

Kris Lovejoy, Global Head of Strategy at Kyndryl, on Chain of Thought ep 62

An endpoint threat model turns a model purchase into a short table: attack, what leaks or breaks, who owns the control. Training teams own data and weights; you own the HTTP contract, quotas, and what you log.

A sketch of the table as code:

# Sketch: one row per attack class for your inference route
THREATS = [
    ("backdoor", "trigger in prod", "provenance + pre-prod red-team"),
    ("evasion", "bypass at runtime", "input limits + secondary policy check"),
    ("inversion", "reconstruct training", "labels only, no logprobs on public API"),
    ("membership", "was record in training?", "coarse outputs + query budgets"),
    ("model_dos", "cost / availability", "rate limits + agent step/token caps"),
]

def review_endpoint(route: str) -> list[tuple[str, str, str]]:
    return [(route, *row) for row in THREATS]
Enter fullscreen mode Exit fullscreen mode

Checklist

  • Inventory every route that calls a model, including nested agent tool loops and the main chat API.
  • For each route, write one line per attack class: asset, attacker goal, your control, owner.
  • Strip public responses down to the minimum fields product needs; put rich telemetry on internal paths only.
  • Put token, step, and spend caps in the gateway rather than in prompt instructions.
  • Re-run the table when you swap weights, add an adapter, or expose a new tool to the agent.

More on this topic, with the related episodes, is on Chain of Thought. It draws on this episode.

Subscribe to the Chain of Thought newsletter for new episodes and write-ups like this one.

Drafted with AI assistance from the episode transcripts.

Top comments (0)