A conversational agent first has to decide what the customer wants. I measured that step for customer-service messages in Brazilian Portuguese (PT-BR), using two small "decider" models and two embedding models. Then I translated the whole dataset into English and ran everything again.
TL;DR
- Decider vs Laya. Off the shelf, the Strands Decider (2B) gets 84.9–85.8% top-1 accuracy, against 50.9–51.8% for Laya multilingual. Its accuracy does not depend on how the options are ordered.
- Embeddings alone win. Qwen3-Embedding-4B similarity, with no decider at all, reaches 94.7%, and no decider improves on it.
- The Portuguese tax. Every component is less accurate in PT-BR than on the same phrases in English. The 0.6B embeddings lose 12 points (p = 0.0014), and the full Decider pipeline loses 8.4 points.
- The checkpoint's language is the biggest lever. Laya's English checkpoint beats its multilingual sibling by 24 points on English text. Laya has no Portuguese checkpoint, which makes a case for PT-BR fine-tuning.
Figure 1. Top-1 accuracy in PT-BR (blue) and in a line-by-line English translation (grey). Bold differences are significant in every seed (exact McNemar test).
The problem: picking the customer's goal
Goal-oriented agents route each message to one of many goals ("pay a bill", "block a card", "talk to a human"). Each goal is defined by a one-line description and a few example phrases. The decision gates everything after it:
- High confidence: the agent acts.
- Middle band: it asks for confirmation.
- Low confidence: it asks a clarifying question.
So you want two things from the classifier: accuracy, and probabilities you can trust.
A new class of small models is built exactly for this. Laya and the Strands Decider answer typed questions ("pick one of these N options") and return a probability per option from a dedicated head, instead of generating text. Most of their published evaluations are in English. I wanted to know how they behave in Portuguese.
How the benchmark works
The setup mirrors a production pipeline:
- Shortlist. An embedding model (Qwen3-Embedding 0.6B or 4B) scores the message against each goal's examples. The 10 most similar goals become the options.
-
Decide. The decider gets a
choicequestion with those 10 goal descriptions and returns a probability for each. - Combine. The two signals are mixed and calibrated:
With a = 0 you trust only the embeddings; with b = 0, only the decider. The weights (a, b, T) are fitted by 5-fold cross-validation on the goals' training examples, never on the test phrases.
The dataset:
- 30 goals from banking, telecom and retail;
- 10 training examples per goal;
- 150 held-out test phrases;
- a line-by-line English translation (same goals, same order, same phrase positions), so every comparison across languages is paired phrase by phrase.
The trap: option order
In production, you naturally present the shortlist sorted by similarity. That means the right answer is usually the first option. A decider that simply prefers the first option would then look exactly as good as your embeddings, without reading the message at all.
So every configuration runs three ways:
- random order, with 3 seeds (the main results);
- similarity order (what production does);
- reversed order (the right answer tends to be last).
That control mattered. Laya multilingual picks the first option in 62% of similarity-ordered questions and loses 11 points when the order is reversed (p = 0.0015). The Strands Decider barely notices: only 4 of 150 phrases change outcome.
Results in Brazilian Portuguese
Random option order, mean ± SD over 3 seeds:
| 0.6B embeddings | 4B embeddings | |
|---|---|---|
| Embeddings alone | 81.3% | 94.7% |
| Laya multilingual alone | 51.8 ± 2.5% | 50.9 ± 1.5% |
| Strands Decider alone | 84.9 ± 1.4% | 85.8 ± 1.4% |
| Strands Decider + embeddings (fitted) | 85.8 ± 0.4% | 89.6 ± 0.8% |
| Decider p95 latency (RTX 5080) | 135 ms | 173 ms |
| Laya p95 latency | 34 ms | 41 ms |
What stands out:
-
Laya is below the embeddings it is supposed to complement. Forcing it into the combination costs 23–31 points; the fit only works if it can set
a = 0and ignore Laya. - The Decider helps a weak retriever, not a strong one. With 0.6B embeddings it adds about 4.5 points (consistent across seeds, not significant at n = 150). With 4B embeddings, the embeddings alone are significantly better than the Decider alone.
- Cross-validated fitting can fool you. With 4B embeddings, fitting on the examples always chose to use the Decider, and lost about 5 points against ignoring it. Inside cross-validation each example is compared against 8 examples per goal instead of 10. That weakens the similarity signal and makes the decider look more useful than it is.
The cost of language
Same models, same phrases, translated into English (paired McNemar tests):
| Component | PT-BR | English | Δ |
|---|---|---|---|
| Embeddings alone, 0.6B | 81.3% | 93.3% | +12.0 (p = 0.0014) |
| Embeddings alone, 4B | 94.7% | 98.0% | +3.3 (n.s.) |
| Strands Decider + 0.6B embeddings | 85.8% | 94.2% | +8.4 (p ≤ 0.004) |
| Strands Decider + 4B embeddings | 89.6% | 95.3% | +5.7 (p ≤ 0.039) |
| Laya multilingual alone | 51.8% | 57.6% | +5.8 |
| Laya English checkpoint alone | — | 81.6% | +24 over multilingual |
The last row is the one I keep coming back to. Laya's English and multilingual checkpoints come from the same release and share an architecture. On English text, training for one language instead of many is worth 24 points, the largest effect in the whole study. For Portuguese, the only option is the multilingual checkpoint, at about 51%.
What this means for your agent
- Spend on the embedding model first. With 10 examples per goal, a good embedding model was the single best component in both languages.
- Test deciders with shuffled options. Production order silently rewards models that just like the first option.
- Tune the combination on held-out phrases, not on your goals' examples.
- Evaluate in your users' language. An English evaluation overstated PT-BR accuracy for every component.
- If you use a decider in PT-BR, plan to fine-tune it. The 24-point gap between Laya's checkpoints shows how much a language-specific model can recover. That is my next experiment.
Reproduce it
Everything is open: the code, both datasets, the server images for both deciders, and the raw per-phrase logs of all 50 runs. You can recompute every number above without a GPU or any model server:
git clone https://github.com/fuljorge/goal-classification-bench
cd goal-classification-bench
uv sync
uv run goalbench report results/v0.1.0/*__* results/v0.2.0/en/*__* --cross-dataset
To run new deciders or models, the servers/ folder has Docker images and plain-venv instructions for both deciders, and goalbench run takes any OpenAI-compatible embeddings endpoint.
fuljorge
/
goal-classification-bench
Reproducible benchmark of choice-style deciders (Laya, Strands Decider) for customer-goal classification in Brazilian Portuguese
goalbench
A reproducible benchmark for customer-goal classification in Brazilian Portuguese with
small "decider" models that answer typed choice questions:
-
Laya (
laya-multilingual, mmBERT-base encoder); -
Strands Decider
(
strands-decider-2B-hobson-v21, Qwen3.5-2B base with LoRA and a pointer head).
Both are compared inside the same two-stage pipeline used by goal-oriented conversational agents:
- an embedding model shortlists the 10 goals most similar to the message;
- the decider picks one of them;
- the two signals are combined and calibrated.
Every call is logged per phrase, so all conditions and statistics are recomputed offline.
-
Method:
docs/method.md. -
Dataset card:
data/README.md. -
Results:
- Brazilian Portuguese:
docs/results.md, with raw logs inresults/v0.1.0/; - language effect, PT-BR vs a parallel English translation
docs/results-language.md, with raw logs inresults/v0.2.0/.
- Brazilian Portuguese:
Results at a glance (v0.1.0)
Top-1 accuracy on 150 held-out phrases, 10-goal shortlist, random option order, mean ± SD over 3 seeds. Latency is the p95…
Caveats
- Synthetic data. The dataset was written by a language model, so absolute numbers are likely optimistic. The paired comparisons are more robust, because every system sees the same phrases.
- Translated English. The English set is a translation, so part of the Portuguese gap may come from translation regularity rather than the language itself.
- Small n. 150 test phrases resolve differences of roughly 8 points or more. Smaller gaps here are consistent but not significant.
The full write-up, with the methodology, all paired tests and threats to validity, is archived with the code on Zenodo: doi:10.5281/zenodo.23233352.
Next in this series: fine-tuning a decider for Brazilian Portuguese, and whether it closes the gap.

Top comments (1)
Two things I'd split before reading the "Portuguese tax" as a decider effect.
First, the pipeline can't beat its shortlist. If recall@10 from the embedding step is lower in PT-BR than in English, part of the 8.4 and 5.7 point drops is phrases whose goal never reached the decider. Reporting recall@10 per language and per embedding size, and then decider accuracy on only the phrases where the right goal was in the ten, would show how much is retrieval and how much is the decider.
Second, the English set is a translation of the PT-BR phrases, so it inherits their intent boundaries and probably their regularity. A cheap check on the "language" reading: back-translate a sample of the English phrases to Portuguese and re-score the embeddings. If the back-translated set lands near the original 81.3% for the 0.6B model, the gap is language. If it lands near 93%, translation smoothing is doing part of the work.
On sizes: 12.0 points on 150 paired phrases fits your p = 0.0014, but the 4B gap of 3.3 points is 5 phrases. I'd quote it as "not distinguishable from zero" rather than a smaller tax, since an exact interval on a handful of discordant pairs spans both zero and the 12 point effect.