Claude Sonnet redacted 20 logs and messages exactly, with or without help. Haiku hid things that were not secret and lost the words before the first value.
A strong model can, and needs no help for it. Claude Sonnet and Claude Haiku got twenty short texts to clean before they went into a log, each holding emails, phone numbers, bank and card numbers, ID numbers, IP addresses or secrets, and asked for the same text back with each value swapped for a fixed placeholder such as [EMAIL]. We compared every answer with the expected text character by character. Sonnet was exact on all twenty, with our instructions or without them. Haiku was exact on eleven without them and fifteen with them, and its most common mistake had nothing to do with secrets: it lost the words that came before the first value it replaced.
Read the full report on AISkills402: https://aiskills402.com/blog/llm-redact-pii-small-model
Top comments (0)