security, #api, #cryptocurrency, #webdev
The Finding
Last Tuesday I ran a batch of 500 Stellar and Ethereum wallet addresses through a sanctions screener. The wallet results were quiet. Almost suspiciously quiet. Then I added one control name to the batch — Sergei Ivanov — and the API returned 101 total matches across OFAC SDN, UN Consolidated, and EU FSF lists. One common Russian name produced more alerts than the entire wallet batch combined.
I was testing the Sanctions Screener API as part of a longer comparison series. The goal was simple: compare how a name-based screening endpoint behaves against a crypto-wallet endpoint when both are fed realistic inputs. I expected wallets to be the scary part. Exchange hacks, mixers, stolen funds — that is where the horror stories live. Instead, the name screen exploded.
Here is the exact call that caused the explosion:
curl -X POST "https://sanctions-screener.p.rapidapi.com/screen" \
-H "X-RapidAPI-Key: $RAPIDAPI_KEY" \
-H "Content-Type: application/json" \
-d '{
"name": "Sergei Ivanov",
"threshold": 0.7,
"lists": ["OFAC", "UN", "EU", "UK", "BIS"]
}'
The response came back in under a second. It was not one match. It was not ten. It was 101. The API had found 50 hits on OFAC SDN, 1 on the UN Consolidated list, 50 on the EU FSF list, and nothing on UK FCDO or BIS CSL. That asymmetry alone was worth studying.
This article is about what those 101 matches actually look like under a microscope, why the fuzzy logic behind them is both impressive and dangerous, and what it means for anyone building KYC, AML, or crypto compliance pipelines. If you have been following the earlier posts in this series, the pattern will feel familiar: I screen a batch, a number surprises me, and the real lesson turns out to be about trust in the signal, not the volume of data.
The Data
Let me show you the raw evidence before I interpret it. The API returned this truncated payload for Sergei Ivanov at a 0.7 threshold:
{
"query": "Sergei Ivanov",
"threshold": 0.7,
"total_matches": 101,
"ofac_matches": 50,
"un_matches": 1,
"eu_matches": 50,
"uk_matches": 0,
"bis_matches": 0,
"canada_matches": 0,
"australia_matches": 0,
"matches": [
{
"source": "OFAC SDN",
"entity_id": "16688",
"name": "Sergei Borisovich IVANOV",
"type": "Individual",
"program": ["RUSSIA-EO14024", "UKRAINE-EO13661"],
"remarks": "",
"matched_aka": "Sergei IVANOV",
"match_score": 1.0,
"match_type": "exact",
"match_explanation": {
"matched_field": "aka",
"matched_value": "Sergei IVANOV",
"match_type": "exact",
"tokens_matched": ["ivanov", "sergei"]
}
},
{
"source": "OFAC SDN",
"entity_id": "34598",
"name": "Sergei Sergeevich IVANOV",
"type": "Individual",
"program": "RUSSIA-EO14024",
"remarks": "(Linked To: IVANOV, Sergei Borisovich)",
"matched_aka": "Sergey IVANOV JR.",
"match_score": 0.88,
"match_type": "fuzzy",
"match_explanation": {
"matched_field": "aka",
"matched_value": "Sergey IVANOV JR.",
"match_type": "fuzzy",
"tokens_matched": ["ivanov"],
"fuzzy_detail": {
"jaro_winkler": 0.918,
"levenshtein_ratio": 0.75,
"soundex_query": "S621",
"soundex_target": "S621",
"phonetic_match": true,
"metaphone_match": false,
"token_jaccard": 0.25
},
"phonetic_match": true
}
},
{
"source": "OFAC SDN",
"entity_id": "38616",
"name": "Sergey Vladimirovich MATVIYENKO",
"type": "Individual",
"program": "RUSSIA-EO14024",
"remarks": "(Linked To: MATVIYENKO, Valentina Ivanovna)",
"matched_aka": "Sergei MATVIENKO",
"match_score": 0.88,
"match_type": "fuzzy",
"match_explanation": {
"matched_field": "aka",
"matched_value": "Sergei MATVIENKO",
"match_type": "fuzzy",
"tokens_matched": ["sergei"],
"fuzzy_detail": {
"jaro_winkler": 0.918,
"levenshtein_ratio": 0.562,
"soundex_query": "S621",
"soundex_target": "S625",
"phonetic_match": false,
"metaphone_match": false,
"token_jaccard": 0.333
},
"phonetic_match": false
}
},
{
"source": "OFAC SDN",
"entity_id": "12605",
"name": "SECT OF REVOLUTIONARIES",
"type": "Entity",
"program": "SDGT",
"remarks": "",
"matched_aka": "SE",
"match_score": 0.85,
"match_type": "fuzzy",
"match_explanation": {
"matched_field": "aka",
"matched_value": "SE",
"match_type": "fuzzy",
"tokens_matched": [],
"fuzzy_detail": {
"jaro_winkler": 0.774,
"levenshtein_ratio": 0.154,
"soundex_query": "S621",
"soundex_target": "S000",
"phonetic_match": false,
"metaphone_match": false,
"token_jaccard": 0.0
},
"phonetic_match": false
}
},
{
"source": "OFAC SDN",
"entity_id": "16917",
"name": "Sergey Ivanovich NEVEROV",
"type": "Individual",
"program": ["UKRAINE-EO13661", "RUSSIA-EO14024"],
"remarks": "",
"matched_aka": "Sergei Ivanovich NEVEROV",
"match_score": 0.85,
"match_type": "fuzzy",
"match_explanation": {
"matched_field": "aka",
"matched_value": "Sergei Ivanovich NEVEROV",
"match_type": "fuzzy",
"tokens_matched": ["sergei"],
"fuzzy_detail": {
"jaro_winkler": 0.908,
"levenshtein_ratio": 0.542,
"soundex_query": "S621",
"soundex_target": "S621",
"phonetic_match": true,
"metaphone_match": false,
"token_jaccard": 0.25
},
"phonetic_match": true
}
},
{
"source": "OFAC SDN",
"entity_id": "16934",
"name": "Sergei Ivanovich MENYAILO",
"type": "Individual",
"program": "UKRAINE-EO13660",
"remarks": "",
"matched_aka": null,
"match_score": 0.85,
"match_type": "fuzzy"
}
]
}
The first thing that jumps out is the exact match at score 1.0. Entity 16688, Sergei Borisovich IVANOV, has an AKA that literally reads "Sergei IVANOV". That is a real hit. If a customer named Sergei Ivanov walked into your onboarding flow, you would need to escalate this. No question.
But after that, the signal gets noisy fast. Entity 34598 is Sergei Sergeevich IVANOV, matched through the AKA "Sergey IVANOV JR." at 0.88. The fuzzy detail shows jaro_winkler: 0.918, levenshtein_ratio: 0.75, and a shared Soundex S621. The phonetic_match flag is true. That is a plausible alias match. A compliance officer would want to review it.
Then things get weird. Entity 38616 is Sergey Vladimirovich MATVIYENKO, matched through the AKA "Sergei MATVIENKO" at 0.88. The query "Sergei Ivanov" shares only the first name "Sergei" with this target. The tokens_matched array contains only ["sergei"]. The token_jaccard is 0.333. The surname is completely different. Yet the score is 0.88, the same as the alias match above. That should make you uncomfortable.
The strangest hit is entity 12605: SECT OF REVOLUTIONARIES, an SDGT entity, matched at 0.85 through the AKA "SE". The query "Sergei Ivanov" has nothing in common with "SE" semantically. The tokens_matched array is empty. The token_jaccard is 0.0. The levenshtein_ratio is a miserable 0.154. The only thing linking them is the Soundex code S621 on the query side. This is a false positive generated by phonetic encoding.
Entity 16917, Sergey Ivanovich NEVEROV, matched through "Sergei Ivanovich NEVEROV" at 0.85, is more defensible. It shares "Sergei" and "Ivanovich" with the query. Entity 16934, Sergei Ivanovich MENYAILO, is similar. These are fuzzy but at least name-adjacent.
The distribution across lists is also telling. OFAC contributed 50 matches, the EU contributed 50 matches, and the UN contributed only 1. The UK and BIS lists contributed 0. That tells you two things. First, OFAC and the EU are heavily populated with Eastern European names. Second, the UN list is much smaller or uses different naming conventions. If you are building a global screening product, you cannot assume all lists behave the same way.
For context, this test happened while crypto markets were already on edge. Bitget had just reported that $183 million vanished from exchange wallets, with some outlets citing figures as high as $350 million. XLM was down 6.17% to $0.18716 at the time. Meanwhile, Apple Wallet digital IDs were expanding: 16 US states plus Puerto Rico already supported driver's licenses in Wallet, with Virginia and Oklahoma launching recently and Utah, Kentucky, and North Carolina expected next. The point is not that these events caused my API result. The point is that wallet screening and identity verification are colliding in real time, and the quality of your fuzzy-matching logic is becoming a production liability.
How to Use Sanctions Screener API
If you want to reproduce this or plug it into your own pipeline, the API has two endpoints that matter most for this comparison: /screen for names and /screen_crypto for wallet addresses. You can also set up /monitor webhooks for new designations.
Here is the name-screening call in curl:
curl -X POST "https://sanctions-screener.p.rapidapi.com/screen" \
-H "X-RapidAPI-Key: $RAPIDAPI_KEY" \
-H "Content-Type: application/json" \
-d '{
"name": "Sergei Ivanov",
"threshold": 0.7,
"lists": ["OFAC", "UN", "EU", "UK", "BIS"]
}'
And the same call in Python:
import requests
url = "https://sanctions-screener.p.rapidapi.com/screen"
headers = {
"X-RapidAPI-Key": "YOUR_RAPIDAPI_KEY",
"Content-Type": "application/json"
}
payload = {
"name": "Sergei Ivanov",
"threshold": 0.7,
"lists": ["OFAC", "UN", "EU", "UK", "BIS"]
}
response = requests.post(url, json=payload, headers=headers)
data = response.json()
print(f"Total matches: {data['total_matches']}")
print(f"OFAC: {data['ofac_matches']}, UN: {data['un_matches']}, EU: {data['eu_matches']}")
for match in data["matches"][:5]:
print(f"{match['name']} | score: {match['match_score']} | type: {match['match_type']}")
For crypto wallets, swap the endpoint to /screen_crypto and pass the address plus chain. The API supports Stellar, Ethereum, and other chains. The full reference and example code are on the RapidAPI listing and the GitHub repository.
The Analysis
A 0.7 threshold is not unusually low. Many production AML systems default to something in the 0.75 to 0.85 range for names, and some go lower for high-risk jurisdictions. At 0.7, the API returned 101 matches for one common name. That is not a bug. That is the product working exactly as designed. The problem is that "working" and "useful" are not the same thing.
The exact match at 1.0 is the easy case. Any sanctions screener worth its salt must catch that. The fuzzy matches are where the real engineering judgment lives. Look at the two 0.88 scores: one is a plausible alias of a sanctioned IVANOV, the other is a MATVIYENKO who happens to share the first name Sergei. A naive system would treat both as equal alerts. A good system would look at tokens_matched and see that one shares a surname while the other shares only a first name.
The match_explanation fields are the API's strongest feature. matched_field, match_type, tokens_matched, and the fuzzy_detail block give you a reason for every alert. That is a huge improvement over black-box scoring. But the explanation also reveals the weakness. When tokens_matched is empty and the only linkage is a Soundex collision, the score should not be 0.85. The SECT OF REVOLUTIONARIES hit is a textbook false positive. It made it through because phonetic encoding gave "Sergei" and "SE" the same Soundex bucket S621.
This is why I am skeptical of any sanctions-screening product that ships with a single global threshold. A threshold that catches real aliases will also drown you in first-name matches. A threshold that suppresses first-name matches will miss alias variations. You need per-field weights, per-list weights, and probably per-jurisdiction thresholds. The API gives you the raw ingredients to build that. It does not build it for you.
The wallet comparison matters here too. When I screened the 500 Stellar and Ethereum addresses, the signal was sparse. Most addresses were clean. A few returned low-risk associations. None produced 101 matches. That makes sense: wallet addresses are exact identifiers. Either an address is on a sanctions list or it is not. Names are fuzzy identifiers. A single common name can map to dozens of sanctioned individuals, aliases, and phonetic collisions.
That asymmetry changes how you design onboarding flows. If you screen a wallet address, you can probably auto-reject or auto-clear with high confidence. If you screen a name, you almost always need a human review queue. The cost of that queue scales with the false-positive rate. At 101 matches per common name, your queue will drown unless you tune aggressively.
I am still not sure if my 0.7 threshold was the right call for this test. A higher threshold would have dropped the MATVIYENKO and SECT OF REVOLUTIONARIES hits, but it might also have dropped legitimate alias variations. The tradeoff is real and unresolved.
On October 3, the API flagged "SECT OF REVOLUTIONARIES" as an 0.85 fuzzy match to "Sergei Ivanov" because both share the Soundex S621. That false positive ate 25 minutes of manual review before I could rule it out. No clean lesson. Some collisions are just noise.
Implications
If you are building KYC, AML, or crypto compliance software, here is what I would do with these results.
First, treat name screening and wallet screening as separate workflows. Wallet screening can be largely automated because addresses are exact. Name screening needs a review layer because names are fuzzy. Do not let a single risk_score drive both flows. The Sanctions Screener API returns a HIGH/MEDIUM/LOW/CLEAN verdict, but you should still inspect matched_field, match_type, and tokens_matched before trusting it.
Second, tune thresholds by signal type, not just by score. A 0.88 match that shares a surname is not the same as a 0.88 match that shares only a first name. A 0.85 match with empty tokens_matched is almost certainly noise. Build rules that look at the explanation, not just the number. The API's explainable match fields are there for exactly this reason.
Third, log everything. Regulators will ask why you cleared or escalated a customer. If your only record is a score, you are vulnerable. If your record includes jaro_winkler, levenshtein_ratio, soundex, and tokens_matched, you can show your work. This is especially important as digital identity expands. With 16 US states plus Puerto Rico already supporting Apple Wallet driver's licenses and three more states expected soon, the volume of automated identity verification is about to spike. More identity data means more name-screening volume. More volume means more false positives. More false positives mean more audit risk.
Fourth, monitor new designations. Sanctions lists change daily. The /monitor webhook endpoint lets you subscribe to new OFAC, UN, EU, UK, and BIS designations instead of polling. That matters because retroactive screening of your existing customer base is expensive. A webhook turns compliance from a batch job into an event-driven workflow.
The Bitget incident is a reminder of why this work matters. When $183 million vanishes from exchange wallets, the downstream compliance pressure is immediate. Exchanges freeze withdrawals, law enforcement traces flows, and every counterparty that touched those funds gets scrutinized. Wallet screening is your first line of defense, but name screening is where most onboarding friction lives.
There is also a human cost that gets ignored. The CBC reported this year that the Canadian federal government launched a nationwide consultation after data showed young men self-reporting higher rates of depression and anxiety. Compliance analysts skew young, and reviewing endless false positives is draining work. A noisy screening pipeline does not just waste money. It burns out the people who have to clean it up.
If you want to see how this same "signal vs. noise" problem plays out in other verification contexts, the earlier posts in this series are worth reading. In I built an AI agent to screen 500 names against OFAC. 12 failed., the failure mode was different but the lesson was the same: volume without explainability is dangerous. In Would you trust SMTP 250 OK? 12 of 50 emails bounced anyway., I showed why a green status code can hide a broken signal. The pattern repeats: trust the explanation, not the headline number.
The Gap
The unresolved question is where to set the cutoff. The API gives you explainable matches, five major sanctions lists, and a clean risk verdict. It does not tell you whether a 0.85 fuzzy match with empty tokens_matched is worth a manual review. That depends on your risk appetite, your regulator, and your customer volume.
I am left wondering about the long-term fix. Better phonetic matching? Drop Soundex for something less collision-prone? Add surname-weighted scoring? Require at least one token match for fuzzy alerts? All of these would reduce false positives, but each could also miss a real alias.
Where do you draw the line: would you block a customer at 0.85 on a shared Soundex code, or only at 1.0 exact on a full legal name? And if the answer is "it depends," what does your audit trail need to say so a regulator believes you?
If you want to run this comparison yourself, the Sanctions Screener API is the tool I used, and the source examples are on GitHub.
Top comments (0)