A few months ago I spent a couple of days debugging blocking mismatches that I couldn't explain. I'd added a Japanese securities data source and a bunch of subsidiary nodes weren't matching their Korean counterparts. Asked for help checking the blocking key function. Got back a version that looked right. Ran it. Still had mismatches. Tried again.
Eventually I just pulled up the raw input data and looked at it. The source had a lot of entries formatted as 合同会社 XYZ. That's a godo kaisha — basically a Japanese LLC. My blocking key function stripped 株式会社 and a few other common forms, but I'd never added 合同会社. So those records were producing blocking keys that didn't match their Korean counterparts, and those pairs never entered the candidate pool.
The fix was one regex line. I'd lost two days on it.
What I kept thinking about after was: the AI couldn't have caught this. It was generating correct general blocking functions. But whether 合同会社 needed to be in the strip list depended entirely on what was in my specific dataset. You can't answer that from the code. You have to look at the data.
Since then I've started paying attention to when a question I'm asking is really a "what's in my data" question in disguise. Confidence threshold questions often are. I'll ask what threshold I should use and get an answer that's been reasoned out from general principles. It'll sound right. But the right threshold for me depends on what the Korean-English pairs in my training set actually look like, and what it costs me when the resolver merges two nodes that shouldn't be merged versus when it misses a merge it should have made. That's not in any prompt I can write.
I haven't fully figured out how to screen for these. Sometimes I catch it before I go down a path, sometimes I don't until I've already spent time on the wrong thing. The 合同会社 case was the one where I lost the most time to it.
I build er-api, a multilingual entity resolution service for Korean, Japanese, Chinese, and English corporate data. More at hannune.ai.
Top comments (2)
The corporate suffix case hits close to home. Models generate clean normalization logic against textbook examples, but they have zero intuition for the messy edge formats lurking in regional data dumps. If ten percent of an ingested table uses an unhandled corporate prefix like 合同会社, the blocking function silently splits candidate pairs before the resolver even runs, and unit tests will not catch it unless you already know the anomaly exists.
The asymmetry in threshold tuning is the other blind spot. A generic prompt assumes balanced loss, but in graph pipelines a false merge poisons every downstream traversal permanently, while a false split is just a missed link you can review later. General reasoning cannot infer that operational blast radius without seeing the production wreckage first.
Reid, that's the framing I've started using too. When I explain it to stakeholders the terminology that actually lands is "the merge is permanent and the split isn't" rather than blast radius language. People get it immediately once you put it that way. The unhandled suffix issue follows me across Korean data too: 협회 and 사단법인 show up constantly in specialty registries and are completely absent from the kind of training sets where normalization logic gets benchmarked, so you only find them when you look at the actual ingested tables.