<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Yaoshen Luo</title>
    <description>The latest articles on DEV Community by Yaoshen Luo (@yaoshen_luo_9d969ebd998fc).</description>
    <link>https://dev.to/yaoshen_luo_9d969ebd998fc</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1529531%2F5af1371a-c5dd-4bf0-952b-344a89930aaa.jpg</url>
      <title>DEV Community: Yaoshen Luo</title>
      <link>https://dev.to/yaoshen_luo_9d969ebd998fc</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/yaoshen_luo_9d969ebd998fc"/>
    <language>en</language>
    <item>
      <title>Hello dev.to! Here is my BenchmarkingRealWork series case 01.</title>
      <dc:creator>Yaoshen Luo</dc:creator>
      <pubDate>Tue, 29 Sep 2026 11:34:38 +0000</pubDate>
      <link>https://dev.to/yaoshen_luo_9d969ebd998fc/hello-devto-here-is-my-benchmarkingrealwork-series-case-01-378p</link>
      <guid>https://dev.to/yaoshen_luo_9d969ebd998fc/hello-devto-here-is-my-benchmarkingrealwork-series-case-01-378p</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/yaoshen_luo_9d969ebd998fc/benchmarking-real-work-case-01-how-i-built-a-voice-agent-benchmark-from-real-customer-failures-12p" class="crayons-story__hidden-navigation-link"&gt;Benchmarking Real Work - Case 01: How I Built a Voice Agent Benchmark from Real Customer Failures&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/yaoshen_luo_9d969ebd998fc" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1529531%2F5af1371a-c5dd-4bf0-952b-344a89930aaa.jpg" alt="yaoshen_luo_9d969ebd998fc profile" class="crayons-avatar__image"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/yaoshen_luo_9d969ebd998fc" class="crayons-story__secondary fw-medium m:hidden"&gt;
              Yaoshen Luo
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                Yaoshen Luo
                
                
              
              &lt;div id="story-author-preview-content-4771088" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/yaoshen_luo_9d969ebd998fc" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1529531%2F5af1371a-c5dd-4bf0-952b-344a89930aaa.jpg" class="crayons-avatar__image" alt=""&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;Yaoshen Luo&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/yaoshen_luo_9d969ebd998fc/benchmarking-real-work-case-01-how-i-built-a-voice-agent-benchmark-from-real-customer-failures-12p" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Sep 29&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/yaoshen_luo_9d969ebd998fc/benchmarking-real-work-case-01-how-i-built-a-voice-agent-benchmark-from-real-customer-failures-12p" id="article-link-4771088"&gt;
          Benchmarking Real Work - Case 01: How I Built a Voice Agent Benchmark from Real Customer Failures
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag crayons-tag--filled  " href="/t/discuss"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;discuss&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/ai"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;ai&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/agentaichallenge"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;agentaichallenge&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/agents"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;agents&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
          &lt;a href="https://dev.to/yaoshen_luo_9d969ebd998fc/benchmarking-real-work-case-01-how-i-built-a-voice-agent-benchmark-from-real-customer-failures-12p" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left"&gt;
            &lt;div class="multiple_reactions_aggregate"&gt;
              &lt;span class="multiple_reactions_icons_container"&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/raised-hands-74b2099fd66a39f2d7eed9305ee0f4553df0eb7b4f11b01b6b1b499973048fe5.svg" width="18" height="18"&gt;
                  &lt;/span&gt;
              &lt;/span&gt;
              &lt;span class="aggregate_reactions_counter"&gt;1&lt;span class="hidden s:inline"&gt;&amp;nbsp;reaction&lt;/span&gt;&lt;/span&gt;
            &lt;/div&gt;
          &lt;/a&gt;
            &lt;a href="https://dev.to/yaoshen_luo_9d969ebd998fc/benchmarking-real-work-case-01-how-i-built-a-voice-agent-benchmark-from-real-customer-failures-12p#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              &lt;span class="hidden s:inline"&gt;Add&amp;nbsp;Comment&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            5 min read
          &lt;/small&gt;
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


</description>
      <category>agents</category>
      <category>ai</category>
      <category>showdev</category>
      <category>testing</category>
    </item>
    <item>
      <title>Benchmarking Real Work - Case 01: How I Built a Voice Agent Benchmark from Real Customer Failures</title>
      <dc:creator>Yaoshen Luo</dc:creator>
      <pubDate>Tue, 29 Sep 2026 11:31:45 +0000</pubDate>
      <link>https://dev.to/yaoshen_luo_9d969ebd998fc/benchmarking-real-work-case-01-how-i-built-a-voice-agent-benchmark-from-real-customer-failures-12p</link>
      <guid>https://dev.to/yaoshen_luo_9d969ebd998fc/benchmarking-real-work-case-01-how-i-built-a-voice-agent-benchmark-from-real-customer-failures-12p</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Customer complaints are not the problem definition; they are the signal.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This post captures a real-world case of agent benchmarking: how I built v1 of our benchmark from real customer failure logs—without prior domain expertise in voiceprint recognition—and used it to drastically improve the production experience of a Voice Agent.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Complaints Are Real Signals, but They Aren’t Actionable
&lt;/h2&gt;

&lt;p&gt;I took over a product module that had strong customer demand, but consistently negative user sentiment.&lt;/p&gt;

&lt;p&gt;For any voice-first agent, voiceprint recognition (speaker identification) is foundational. In multi-party conversations, whether an agent can correctly identify who is speaking significantly affects the conversation flow and determines whether conversational memory gets attributed to the right user.&lt;/p&gt;

&lt;p&gt;When I stepped in, I had almost zero context.&lt;/p&gt;

&lt;p&gt;All I had were escalating customer inquiries demanding optimization progress, paired with frustrated feedback. But the input was always along the lines of: "Hey, we hit another misidentification issue earlier." In practice, this kind of feedback is completely non-actionable.&lt;/p&gt;

&lt;p&gt;An interaction with an agent is composed of turn-by-turn dialogue. Customers can rarely pinpoint exactly which turn was misattributed, or when the system simply failed to detect a speaker altogether.&lt;/p&gt;

&lt;p&gt;Still, customer frustration is always a solid lead.&lt;/p&gt;

&lt;p&gt;I asked them to run another test round and send over the failure cases they ran into. Sifting through those fragmented logs and session artifacts, the problem's shape started to emerge.&lt;/p&gt;

&lt;p&gt;Root-causing speaker ID failures on a per-incident basis was painfully inefficient. Every single dialogue turn can trigger a speaker prediction, and that prediction ripples directly into the language model’s prompt context.&lt;/p&gt;

&lt;p&gt;I had to translate "I hit another identification bug, go look into it" into "In this 30-turn session, turns 10 and 15 misclassified the speaker as someone else, while turn 7 dropped the detection entirely."&lt;/p&gt;

&lt;p&gt;You have to convert fuzzy feedback into granular, structured issues before you can take any engineering action.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Reproduction, Annotation, and Evaluation: Turning Conversations into Repeatable Datasets
&lt;/h2&gt;

&lt;p&gt;I looked around—there were no off-the-shelf internal tools for this.&lt;/p&gt;

&lt;p&gt;This is a classic engineering reality: you have the application code to ship features, but zero tooling to reproduce failures. Which means you have no mechanism to handle post-launch customer escalations or iterate with confidence.&lt;/p&gt;

&lt;p&gt;I treated this lack of full session context capture and replay tooling as problem zero.&lt;/p&gt;

&lt;p&gt;Once we had a replay mechanism in place, the path forward became clear: I could listen to a 10-turn dialogue trace and label the ground-truth speaker identity for each utterance. With labeled ground truth, replaying the problematic session against the engine and comparing actual vs. expected outputs gives you concrete metrics.&lt;/p&gt;

&lt;p&gt;Here, the metrics were clean and straightforward: per-turn accuracy and recall for speaker ID.&lt;/p&gt;

&lt;p&gt;I put together a lightweight web interface for the annotation flow. It sequentially loaded the dialogue context and audio clips so I could tag speakers in real time as I listened through. Hit save, and you have a structured, annotated sample.&lt;/p&gt;

&lt;p&gt;I obsessed over the ergonomics of this annotation tool. If we needed to quickly scale the dataset to cover diverse production scenarios, my teammates and I had to be able to jump in and label without friction.&lt;/p&gt;

&lt;p&gt;With reproduction and labeling solved, I wired up the evaluation runner. Just like that, we closed the loop on our first repeatable evaluation pipeline: replay, annotate, execute, and score.&lt;/p&gt;

&lt;p&gt;I brought in a few colleagues to simulate real-world conversations with the agent, covering single-speaker and multi-speaker turn-taking scenarios. All they had to do was talk naturally to the agent. Leveraging the internal tool, I quickly built out the dataset with hundreds of labeled conversational turns. It came together faster than expected: v1 of the benchmark was live.&lt;/p&gt;

&lt;p&gt;I couldn't wait to run the baseline evaluation. And... it was predictably disappointing. Per-turn identification accuracy was abysmal across the board—nowhere near the open-source model's performance on public benchmarks.&lt;/p&gt;

&lt;p&gt;What was falling apart? Here is the dirty work driven by the benchmark.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. How the Benchmark Drove Root Cause Analysis and Iterations
&lt;/h2&gt;

&lt;p&gt;I’m by no means an algorithm expert in voiceprint recognition; it was completely new territory for me.&lt;/p&gt;

&lt;p&gt;Yet, through the benchmark runs and manual test sweeps, I noticed that speaker ID accuracy correlated heavily with audio sample duration. I sliced the benchmark across several audio chunk durations, and the differences were clear.&lt;/p&gt;

&lt;p&gt;I used an agent to run an academic literature pass and confirmed that the dependency of speaker identification on sampling duration has been thoroughly documented.&lt;/p&gt;

&lt;p&gt;I redesigned the module’s audio ingestion pipeline: specifically when inference triggers after receiving the audio stream, and the duration of the audio sample. The goal was to ensure speaker metadata was injected into the LLM context reliably and promptly, regardless of how long or short the user utterance was.&lt;/p&gt;

&lt;p&gt;Boom! The benchmark scores jumped immediately.&lt;/p&gt;

&lt;p&gt;We weren't out of the woods yet. We had addressed the core identification issue, but the actual conversational experience still had problems.&lt;/p&gt;

&lt;p&gt;Whenever users switched turns in a group setting, the agent’s memory and response phrasing still degraded. The root cause: the speaker metadata was poorly integrated into the prompt layer.&lt;/p&gt;

&lt;p&gt;Once we structured and normalized the voiceprint metadata into the LLM's prompt context, the problem disappeared. Both the v1 benchmark metrics and hands-on testing saw dramatic, measurable improvements.&lt;/p&gt;

&lt;p&gt;The experience of dynamic multi-user turn-taking and discussions quickly won over my colleagues—especially given how painful it had been when they were recording the baseline test data.&lt;/p&gt;

&lt;p&gt;When I invited customers back for another evaluation, many responded positively. They also continued to report cases that fell short of expectations.&lt;/p&gt;

&lt;p&gt;Digging into those new cases pointed the finger back at the benchmark itself: while v1 had driven real performance gains, its sample distribution had systematic bias.&lt;/p&gt;

&lt;p&gt;The utterance length distribution in v1 did not match production traffic. It was skewed heavily toward longer sentences, making the evaluation overly optimistic. Once we re-sampled the benchmark to mirror real-world utterance length distributions, our accuracy numbers dropped back down.&lt;/p&gt;

&lt;p&gt;Only then did we adjust the recognition model.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. From a Static Benchmark to Continuous Evaluation: The Scale and Coverage Dilemma
&lt;/h2&gt;

&lt;p&gt;Following these iterations, some customers who had previously been disappointed with the feature confirmed that the core experience had noticeably improved.&lt;/p&gt;

&lt;p&gt;Building v1 of the benchmark wasn't technically complex, nor did it take massive engineering hours. It didn't try to cover every edge case under the sun, but it gave me what mattered: the ability to reproduce customer failures, pinpoint root causes, and verify whether our patches actually worked.&lt;/p&gt;

&lt;p&gt;To me, this is what defines a viable benchmark: it doesn't need to be exhaustive or complicated out of the gate, but it must change how you understand the problem and unblock your next engineering action.&lt;/p&gt;

&lt;p&gt;That said, v1 is still just a static, repeatable test suite. Turning it into a continuous evaluation system means tackling harder operational questions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;How do we identify failures in real user sessions that are worth adding to the benchmark?&lt;/li&gt;
&lt;li&gt;How do we keep human annotation reliable and consistent while controlling costs?&lt;/li&gt;
&lt;li&gt;How do we continuously expand coverage across hardware devices, new users, variable utterance lengths, and multi-speaker conversations?&lt;/li&gt;
&lt;li&gt;How do we manage benchmark versions and keep them aligned with real user distributions?&lt;/li&gt;
&lt;li&gt;How do we feed newly discovered production issues back into the next evaluation cycle?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These questions draw the line between a static benchmark and a continuous evaluation system.&lt;/p&gt;

&lt;p&gt;V1 helped us reproduce and fix the issues we could see. The next engineering challenge is building a pipeline that continuously uncovers the failures we haven't seen yet.&lt;/p&gt;

&lt;p&gt;I'll dive into scaling trustworthy annotation, maintaining scenario coverage, and building a continuous evaluation loop in the upcoming pieces of the &lt;em&gt;Benchmarking Real Work&lt;/em&gt; series.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agentaichallenge</category>
      <category>agents</category>
      <category>discuss</category>
    </item>
  </channel>
</rss>
