DEV Community

Cover image for I built a harvest tripwire, then almost let an LLM decide who looks legit
Klaus Byskov Pedersen
Klaus Byskov Pedersen

Posted on AI-assisted

I built a harvest tripwire, then almost let an LLM decide who looks legit

Free API tiers get farmed. That is not news. What surprised me was how much the farming looked like real evaluation, and how badly a model did at telling the two apart.

This is a Companydata story: Danish company registry data over HTTP, a free monthly quota, and a tripwire that throttles keys that suddenly look up hundreds of different companies in an hour.

If you run a free or freemium API, I want your war stories in the comments. I am still deciding whether to pool free quota by IP after a burn. How did you solve the "looks like a tester, is actually rotation" problem without nuking real evaluators?

The symptom that felt wrong

My harvest tripwire was doing its job: when a free key tore through a wide set of company lookups in a short window, I auto-flagged it and slowed it to one request per minute. Alerts came into Telegram. Several looked like ordinary builders testing an integration. I vouched a few of them by hand.

That felt expensive. Every false positive is a conversion you might have killed with politeness. So the tempting next step showed up fast: put a model on the alert and ask, "does this look like a legit user?"

I had already been evaluating Jev (TypeSafe System One) in another project. Identity-only signals, a handful of typed questions, cheap tokens. Perfect candidate for a quick legitimacy gate. Or so it seemed.

What the data actually showed

Three "testers" in a couple of days were not three testers. Same egress network, same curl client, same habit of burning the free quota to the exact limit, then opening a fresh account. Real people, real company, still 1,500 free calls in 48 hours by one client.

Quota emails were opened. Pricing was known. A free re-signup was cheaper than a small credit pack. Information was not the bottleneck. Friction on the free path was.

That finding mattered more than any model score. The tripwire was not wrong about the pattern. The human vouching step was wrong about independence of accounts.

I ran Jev anyway

I still ran the experiment. Twenty-seven accounts that already had an abuse field: seven I had vouched as legit, twenty from an earlier abusive ring. Identity-only state into Jev.

Jev found every vouched account when I asked for a hard "legit" verdict. It also waved through several of the ring. Probability scores sat in a mushy band for both classes. No useful threshold. When I added behaviour, Jev flipped most of the vouched accounts to abusive, which, given the rotating-quota story, was closer to the truth and also useless as a gate for "please don't annoy real customers."

A dumb rule score on identity (Google signup, non-freemail domain, name shape, and so on) separated the old ring from clean-looking identities, and would have waved the rotating-quota trio straight through. Identity is context for a human on-call, not a gate.

Decision: Jev stays out of the tripwire. Same conclusion I reached with Jev on another product evaluation earlier that month.

What I shipped instead

Two boring, high-leverage changes.

Related accounts on the admin alert. When a key trips, the Telegram message now lists other accounts that share the client address or the exact display name, with usage and state. Admin group only. Nothing about other users ever goes into a user-facing response. The first live run immediately taught me to exclude my own frontend service identity, which had been logging anonymous page views under the visitor's address and therefore headed every related list.

An honest 429. A flagged key used to get the same "Rate limit exceeded" body as an ordinary minute limit. Scripts backed off and nobody wrote in. Flagged keys now send X-RateLimit-Reason: under-review and a body that says I am reviewing a free key that looked up an unusual number of companies, that nothing is blocked forever, and how to buy credits or mail support. Ordinary minute-limit responses are unchanged.

I deliberately did not auto-withhold a new free quota from an address that already exhausted one this month. That is a product call I want to make after living with the better alerts.

What I would tell someone building the same thing

  1. Harvest detection on breadth and velocity works. The failure mode is the human story you tell yourself about who is behind the key.
  2. Account linking for operators beats clever legitimacy scoring for users.
  3. If you throttle someone, say so in the response. Silent identical 429s train both scrapers and real integrators to shrug.
  4. Try the model if you want the data. I named Jev here because people ask. For this problem it did not earn a place in the critical path.

Your turn

If you have shipped something similar (Stripe-style fingerprinting, shared free pools, device attestation, payment before second account, soft CAPTCHA on signup bursts, or something weirder), drop it below. Especially useful:

  • What signal actually caught rotation without false-flagging agencies and shared offices?
  • Did you ever put an LLM on abuse triage, and did it earn its keep?
  • How do you talk to a real evaluator who tripped the wire without sounding like a bank fraud team?

Companydata is live at companydata.dk. The API docs cover the rate-limit header. If you are evaluating and hit the review throttle for real work, mail support and I clear it by hand.

Top comments (2)

Collapse
 
officialmailkr profile image
오피셜메일 •

일반 분당 제한과 검토 중인 키에 같은 429를 보내던 부분이 특히 와닿네요. IP로 할당량을 묶기 전에는 공유 사무실의 정상 계정을 비교군으로 남겨 두면 좋겠습니다. 지원 요청 후 다시 정상화되는 데 걸린 시간도 함께 보면, 오탐 비용을 사용량과 별도로 확인할 수 있겠어요.ㄴ

Collapse
 
officialmailkr profile image
오피셜메일 •

일반 분당 제한과 검토 중인 키에 같은 429를 보내던 부분이 특히 와닿네요. IP로 할당량을 묶기 전에는 공유 사무실의 정상 계정을 비교군으로 남겨 두면 좋겠습니다. 지원 요청 후 다시 정상화되는 데 걸린 시간도 함께 보면, 오탐 비용을 사용량과 별도로 확인할 수 있겠어요.