Claude Sonnet and Haiku turned ten written specs into test cases, which we ran against 61 planted bugs. The logic held up; the JSON broke more often.
Yes, and the expected values are mostly right. Thirteen tasks over ten small functions, each described by a written spec, went to Claude Sonnet and to Claude Haiku, and both were asked for unit tests in one form: a JSON table of test cases. Then we ran every table against our reference code and against the bugs we had planted in copies of it. Without any help, each model got exactly one expected value wrong across all thirteen tasks. What failed more often was the answer as a file: undefined and NaN written straight into the arguments, or a sentence of explanation before or after the table. A test runner that loops over the file stops at both, so those tables count as failures even when…
Read the full report on AISkills402: https://aiskills402.com/blog/llm-tests-from-spec-json-first
Top comments (0)