DEV Community

Cover image for Benchmarking Real Work Case 2: Why Good Agent Evaluation Wasn’t Enough for Production
Yaoshen Luo
Yaoshen Luo

Posted on Originally published at aiagentbenchmark.com

Benchmarking Real Work Case 2: Why Good Agent Evaluation Wasn’t Enough for Production

"They used it once and refused to touch it again."

When that feedback landed two weeks in, I could barely believe it. An agent validated against a benchmark built on real-world data had still failed to solve the foundation’s workflow problem.

Only in hindsight did I realize: I had evaluated the agent's processing capability, but I never evaluated whether volunteers could actually use it to get their job done.

1. Every House Has a Kitchen—Why Build a Community Canteen?

A friend working at a non-profit foundation once shared a story with me:

A rural village had poured resources into numerous senior welfare initiatives over the years, yet life satisfaction barely moved. What ultimately made a difference was building a communal canteen.

When the village head first heard the project proposal, he was deeply skeptical: Every single elderly household has its own kitchen. Why on earth do we need a village canteen?

Field visits later validated two straightforward reasons. First, cooking had simply become too physically taxing for the seniors—especially firing up a stove in the sweltering heat, which was pure misery. Second, beyond serving hot meals, the canteen at the village entrance created a regular venue for group interaction. It gave them a place to run into people and chat, directly addressing their need for social connection.

This counterintuitive, genuine need was surfaced by foundation staff and volunteers through round after round of authentic home-visit conversations with local seniors.

Their organization runs diverse programs nationwide. These projects rely on cohorts of short-term volunteers who pay regular visits to seniors across different age groups, speaking varied regional dialects. Those conversational touchpoints hold the richest, yet most frequently overlooked, context regarding the seniors' current living conditions and real needs.

While my friend had the foresight to archive these interview audio recordings, data collection still leaned entirely on volunteers manually filling out questionnaires.

That survey workflow was time-consuming and labor-intensive, making standardization difficult. More crucially, longitudinal insights—such as tracking changes in a senior's condition over time—rarely materialized.

Meanwhile, a lean foundation team couldn't realistically audit hours of raw audio by hand.

They needed a way to truly listen to these seniors at scale.

2. From an Audio File to an Auditable Interview Record

I realized the foundation didn’t need raw audio files; they needed reliable, verifiable interview records. That breakdown came down to three concrete requirements.

First: Accurate content transcription. Across these interviews, seniors speaking regional Chinese dialects was the baseline reality. This required a rigorous vendor evaluation for speech-to-text (STT).

I requested a small batch of authentic interview datasets. These were dialect-heavy recordings—each running over 30 minutes, spanning dozens of back-and-forth turns, with standard Mandarin and local dialects constantly interspersed. I benchmarked several speech recognition vendors that supported dialects against these files.

Evaluation was constrained to spot-checks by team members fluent in the respective dialects, paired with holistic context sanity checks. Through this filtering, I locked down our technical vendor and measured average transcription costs on real-world data.

Second: Accurate speaker identification. Listening closely to the tapes revealed messy, real-world conversational dynamics, typically involving multi-party conversations.

Volunteers often conducted visits in small teams. Meanwhile, the senior being interviewed was frequently surrounded by family members or neighbors chiming in. For instance, if an extraction pipeline relied purely on raw transcripts to log the senior's health updates, it could easily misattribute statements made by the spouse.

However, STT vendors supporting broad dialect coverage rarely offered out-of-the-box speaker identification. In our processing pipeline, we therefore had to run speaker identification on the raw audio of each segmented utterance, injecting speaker identity as a critical context layer.

Third: Grounded rationale extraction. The survey questions had to be answered strictly based on the interview dialogue, and many items had direct evidentiary clues scattered across the transcript.

My primary concern here was hallucination in the underlying LLM—specifically, downstream evaluative conclusions generated without factual backing.

I addressed this with a two-layer design: 1) Structurally, the workflow retained a human-in-the-loop review step for volunteers to verify answers. 2) Technically, in addition to outputting the structured form fields, the agent was forced to simultaneously extract the verbatim dialogue snippets serving as evidence.

This second design unlocked a standardized benchmark to score relevance between the model’s outputs and the source transcript, measuring precision and recall on evidence retrieval. While this doesn't offer an absolute algorithmic guarantee against hallucinations, it effectively benchmarked different model providers. It also enabled runtime quality checks on citation integrity in production, providing a reliable proxy for dialogue session health.

Anchored on these three core principles, I set out to hand-label the small sample of real interview data to stand up our benchmark.

Dialect testing was limited to spot-checking rather than turn-by-turn ground-truth matching, so the benchmark baseline assumed clean, pre-segmented utterances. I built a lightweight internal web tool to load this context, play back audio per utterance, and let me label speaker identities.

For the post-interview survey extraction, I created ground-truth annotations of verbatim evidence and corresponding answers across the sample cases. The evidence retrieval was cleanly testable with deterministic metrics, while answer accuracy was evaluated using an LLM-as-a-judge approach as a directional reference.

With this benchmark running, I vibe-coded an end-to-end web prototype and iterated through multiple testing loops. This involved refactoring the data processing pipeline, redesigning the web interface flow, and selecting key vendors—particularly balancing latency, accuracy, and operational cost.

Ultimately, the prototype reached a very usable level on the benchmark. The dataset was small, but I had reasonable confidence in the pipeline's output quality.

3. The Benchmark Looked Great. Why Did Volunteers Refuse to Use It?

The benchmark numbers gave me strong confidence in this build. I recorded a demo video for my friend, who was also impressed with the output. Within two days, he assigned a remote intern and the volunteer operations lead to pilot the tool in an upcoming regional field visit.

For the next few weeks, I repeatedly checked our dashboard, watching trial volunteers sign up and trigger runs. The counters showed tasks completing normally, and I convinced myself that the tool had solved the foundation's interview processing problem.

Two weeks later, my friend came back with the field report: the end-user experience was terrible. After testing it once, the frontline volunteers refused to touch it again.

I have to admit, when that feedback first came in, I found it hard to believe. With an evaluation backed by real-world audio showing strong results, the actual production experience should not have been so poor.

Where did it break?

I first synced with the intern, who insisted the product vision was completely spot-on and genuinely meaningful. We brainstormed a backlog of incremental feature optimizations. Then, pulled away by other priorities over the next few months, I let the project stall.

Later, my friend insisted on pushing this tool through to full deployment for ongoing volunteer operations. I joined one of their internal working sessions. That was when I realized the intern was working remotely; like me, he had limited visibility into the day-to-day work of frontline volunteers.

In that meeting, I finally spoke with the volunteer coordinator who had reported the poor experience. Listening to her describe the workflow, I caught the core friction point causing the backlash: for volunteers, the operational overhead had multiplied, while the perceived efficiency gain over their legacy process was negligible.

I asked her to walk me through their baseline workflow. They already had a standardized SOP for regional field teams: volunteers simply completed their onsite interview, uploaded the audio file, and manually filled out the survey form.

Volunteers executed this entire legacy flow on their phones via WeChat, simply tapping a Feishu web form link provided by headquarters.

My "solution," by contrast, required volunteers to open a third-party speech SaaS on desktop, register an account, pre-import questionnaire prompt templates, upload audio files turn-by-turn, and export the transcribed extraction as a CSV file.

It was an embarrassingly obvious failure mode. In retrospect, I struggled to understand how I missed such glaring workflow friction when the initial feedback came in.

It boils down to one reality: I wasn’t on the ground.

4. An Agent Only Works When It Disappears into the Existing Workflow

That feedback was useful because the source of the problem was clear.

A straightforward idea surfaced: How could I wire this benchmarked pipeline directly into the volunteers' existing workflow? Ideally, the integration would be completely invisible to them, dropping adoption friction to near zero.

I focused on Feishu, their existing workspace tool. Volunteers were already relying on its native file upload and survey forms for automated data collection. Digging into the platform capabilities, I found that Feishu supported custom workflows combining audio ingestion and LLM execution. I could rebuild the pipeline natively on top of their existing automation infrastructure.

I stood up the working prototype over a single weekend. It mirrored the exact audio submission entry point volunteers were already accustomed to, handling dialect transcription, speaker parsing, and grounded Q&A extraction entirely downstream.

The turnaround was immediate and stark.

The intern quickly grasped why our earlier disconnect from the operational context had derailed us. After reviewing the prototype architecture, he took over and built out the complete internal production implementation. He also spearheaded rollout validation across pilot regions, mapping the geographic distribution of elderly dialects to define the hard operating boundaries of the system. In regions speaking dialects beyond our coverage limits, the legacy manual workflow was kept intact.

Only then was the project truly in production use.

Whenever I look back at this case, I stay paranoid about one question: What are the actual measurement boundaries of the benchmark in front of me? And what is it failing to measure?

This experience didn't invalidate our benchmark. What it made painfully clear was that the benchmark answered a very narrow question: Given an audio file, can the agent generate an accurate, auditable interview record? What it completely failed to answer was: Can a volunteer, as an end user, complete their end-to-end task inside their existing workflow with less operational friction?

I now view the workflow itself as part of the evaluation, rather than just scoring raw model outputs. How input enters the system, how outputs are reviewed and submitted, and whether human users continue using the tool all help determine whether an agent delivers sustained real-world value.

Top comments (0)