Two Claude models labelled 25 support messages for us. Category and language came out almost always right; priority tracked the tone instead of the problem.
Not consistently. Sonnet and Haiku, two Claude models, each received 25 short customer messages with a request for three labels in JSON: a category from a fixed list, a priority from urgent, high, normal and low, and a language code. Without written rules for choosing the priority, Sonnet matched our expected labels on 21 messages and Haiku on 16. The category and the language were nearly always right. Twelve of the thirteen misses were about priority, and ten of those twelve went up rather than down: shouting, anger, a request for a quote and a delayed parcel were enough to move it. The two models also disagreed with each other on seven messages. Once our rules were added as a skill, both…
Read the full report on AISkills402: https://aiskills402.com/blog/llm-ticket-triage-tone-priority
Top comments (0)