DEV Community

Aamer Mihaysi
Aamer Mihaysi

Posted on

Prompt search is a hill-climber, and accuracy is the wrong hill

I once shipped a prompt that scored 0.94 on my eval set and was useless in triage. Not wrong, exactly. Just useless — it ranked the one case I needed to see at position nine, behind eight things that were fine.

That's the whole article, really. But the mechanism is worth spelling out, because it wasn't a fluke. It was the thing I asked for.

The myth

"If I optimize my prompts against accuracy, I get a better model."

Every prompt optimizer I've used — DSPy-style compilers, evolutionary search over instruction strings, the OPRO-ish loops people wire up themselves — works the same way underneath. You hand it a scalar. It hill-climbs. It has no idea what the scalar means. Accuracy, F1, BLEU, a hand-rolled score from a judge model: all identical to the search. It finds the easiest high point in the space you defined and defends it.

So the real question isn't "does prompt optimization work". It's "what did you point it at".

Why accuracy is a bad hill on imbalanced data

My last clinical-ish eval set was about 4% positive. Rare finding, lots of normal scans, which is what real clinical data looks like.

On that set, a prompt that answers "nothing abnormal" to everything scores 0.96 accuracy. It is also the single most dangerous output you could ship. And it's the easiest point in the search space to reach — no reasoning, no grounding, no long instruction to get wrong. Your optimizer will find it in a couple of generations and then fight every mutation away from it, because everything else looks worse on the number you gave it.

That's not a bug in the optimizer. That's a correct optimizer solving the problem you wrote down.

The subtler failure is the one that got me. Accuracy is thresholded. It collapses a score into a yes/no and throws the score away. Two prompts can hit identical accuracy with completely different ranking behavior — one separates the positives cleanly at the top, one scatters them through the middle and gets lucky at the cutoff. Your search sees a tie. It picks whichever one it evaluated first, or whichever one is more confidently wrong on the majority class, because confidence often tracks the majority pattern.

If your deployment decision is "review the top k", you just optimized for a metric that doesn't describe your deployment.

Multimodal makes this worse, not better. When the prompt has to coordinate an image encoder, a text instruction and a scoring rubric, the search space is enormous and the degenerate solution — ignore the image, answer the prior — is cheap to find and cheap to keep. Bigger space, same bad hill.

What the paper does about it

This is the part I actually wanted to write about.

The setup: you're running prompt search, and you keep a matrix of results — rows are candidate prompts, columns are eval instances, cells are outcomes. If the cell is "was this instance correct", the column average is accuracy. That's the fitness function. That's the hill.

The paper's move is to change what a row is. Instead of one row per instance, you build rows over positive-negative pairs. The cell becomes "did this prompt score the positive above the negative". Now the column average is AUROC.

Same scores. Same model calls. Different bookkeeping. You didn't buy a better model — you stopped lying to your search loop about what you wanted. The optimizer still hill-climbs; it just climbs a surface that has the same shape as your deployment decision.

That's the kind of trick I like. It costs nothing at inference time, and it only requires that your harness kept raw scores instead of booleans. Which, if you're like me, it didn't.

What I changed in my own harness

Three things, all boring:

  1. Store the raw score, not the pass/fail. Along with the prompt hash and the instance id. If your eval harness only writes booleans, you've already destroyed the ranking signal, and no amount of clever search will get it back. This is the actual fix. Everything else is downstream.

  2. Compute AUROC next to accuracy, always. Not instead of — next to. When they disagree, that disagreement is the interesting part of the run, and it usually means your positive class is too small for accuracy to say anything.

  3. Make the search read the metric that matches the decision. If I ship a threshold, I optimize precision at that operating point. If I ship a ranked list for a human to review, I optimize ranking. Those are different numbers and I've stopped pretending one stands in for the other.

Where I'm not sure

Pairwise rows blow up. Positives times negatives — with 4% positives over a few thousand instances that's manageable; with a genuinely rare class it isn't, and you end up sampling pairs, which means your AUROC estimate carries variance your optimizer will happily exploit. I haven't run this at a scale where that bites yet, so I don't know how ugly it gets.

Ties are the other one. AUROC needs a convention for tied scores — half credit is standard — and discrete prompt outputs tie constantly. If your optimizer is comparing two prompts that both emit "no", the pairwise matrix fills with ties and the signal thins out.

And ranking isn't calibration. A prompt can rank beautifully and still be badly calibrated at the threshold you actually deploy at. AUROC doesn't care. Your on-call engineer does.

The takeaway

Prompt optimization is search. Search is honest — it finds exactly what you asked for, including the degenerate version. The metric your search reads has to be the metric you deploy on, and on imbalanced data, accuracy is almost never that metric.

If you take one thing: go look at whether your eval harness stores scores or booleans. That's the whole fix, and it's free.

Paper: https://arxiv.org/abs/2609.40361v1

Top comments (0)