<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Gorakhnath Yadav</title>
    <description>The latest articles on DEV Community by Gorakhnath Yadav (@gorakh13).</description>
    <link>https://dev.to/gorakh13</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1650053%2Fe62a2d5b-dd1c-4e98-bcde-6f1e80e5f615.jpg</url>
      <title>DEV Community: Gorakhnath Yadav</title>
      <link>https://dev.to/gorakh13</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/gorakh13"/>
    <language>en</language>
    <item>
      <title>New blog</title>
      <dc:creator>Gorakhnath Yadav</dc:creator>
      <pubDate>Thu, 17 Sep 2026 22:03:03 +0000</pubDate>
      <link>https://dev.to/gorakh13/-40gf</link>
      <guid>https://dev.to/gorakh13/-40gf</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/gorakh13/my-paper-reader-answered-questions-for-weeks-i-never-checked-if-it-was-right-57n8" class="crayons-story__hidden-navigation-link"&gt;My Paper Reader Answered Questions for Weeks. I Never Checked If It Was Right.&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/gorakh13" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1650053%2Fe62a2d5b-dd1c-4e98-bcde-6f1e80e5f615.jpg" alt="gorakh13 profile" class="crayons-avatar__image" width="800" height="1062"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/gorakh13" class="crayons-story__secondary fw-medium m:hidden"&gt;
              Gorakhnath Yadav
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                Gorakhnath Yadav
                
                
              
              &lt;div id="story-author-preview-content-4679808" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/gorakh13" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1650053%2Fe62a2d5b-dd1c-4e98-bcde-6f1e80e5f615.jpg" class="crayons-avatar__image" alt="" width="800" height="1062"&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;Gorakhnath Yadav&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/gorakh13/my-paper-reader-answered-questions-for-weeks-i-never-checked-if-it-was-right-57n8" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Sep 17&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/gorakh13/my-paper-reader-answered-questions-for-weeks-i-never-checked-if-it-was-right-57n8" id="article-link-4679808"&gt;
          My Paper Reader Answered Questions for Weeks. I Never Checked If It Was Right.
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/ai"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;ai&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/rag"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;rag&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/llm"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;llm&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/voiceai"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;voiceai&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
            &lt;a href="https://dev.to/gorakh13/my-paper-reader-answered-questions-for-weeks-i-never-checked-if-it-was-right-57n8#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              &lt;span class="hidden s:inline"&gt;Add&amp;nbsp;Comment&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            18 min read
          &lt;/small&gt;
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


</description>
    </item>
    <item>
      <title>My Paper Reader Answered Questions for Weeks. I Never Checked If It Was Right.</title>
      <dc:creator>Gorakhnath Yadav</dc:creator>
      <pubDate>Thu, 17 Sep 2026 21:53:16 +0000</pubDate>
      <link>https://dev.to/gorakh13/my-paper-reader-answered-questions-for-weeks-i-never-checked-if-it-was-right-57n8</link>
      <guid>https://dev.to/gorakh13/my-paper-reader-answered-questions-for-weeks-i-never-checked-if-it-was-right-57n8</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;I've been messing with a small project called Talkit, which reads research papers aloud and answers spoken questions about them. The answering half is retrieval-augmented generation, and this is what it took to find out whether it worked: grounding, generation, evals, tracing.&lt;/li&gt;
&lt;li&gt;There's no vector store. A paper fits in context, so the retrieval step is the whole paper plus the passage that was just read out.&lt;/li&gt;
&lt;li&gt;Evals run through DeepEval against ten questions about the Attention paper, scored four ways. Two of the questions have no answer in the paper.&lt;/li&gt;
&lt;li&gt;The app already emitted OpenTelemetry traces, so pointing them at Confident AI was four environment variables and one small code change. Making a span &lt;em&gt;scoreable&lt;/em&gt; took a dozen attributes, and they stay off by default because they carry the question and the paper text.&lt;/li&gt;
&lt;li&gt;On the first real run, every answer scored 1.00 on faithfulness. The failures were elsewhere: a model that sometimes never answers, and one of my own goldens asking for more than its question did.&lt;/li&gt;
&lt;li&gt;The eval suite isn't in CI. It costs credits on every run, and the model under test sometimes returns no answer at all.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I've been chipping away at a small project called Talkit. It reads research papers aloud and lets you interrupt it. You say "wait, why do they scale that?", the narration stops, it answers in two or three spoken sentences, and then it picks up where it left off. That part works. I use it.&lt;/p&gt;

&lt;p&gt;What I could not have told you, for weeks, was whether the answers were any good.&lt;/p&gt;

&lt;p&gt;The evidence I had was two scripts. One uploads the Attention paper, asks a single question, and checks that &lt;em&gt;an&lt;/em&gt; answer comes back. The other does the same in Hinglish. An answer that invented a BLEU score would have passed both. So would an answer that confidently explained the wrong paragraph. The endpoint returned 200 and the suite stayed green, and a green suite quietly becomes the thing you trust.&lt;/p&gt;

&lt;p&gt;The other reason I put it off is that the RAG advice I kept reading didn't fit. Every guide opens with chunking and embeddings and retrieval metrics, and Talkit has no vector store at all: the paper fits in the model's context, so it sends the whole thing, which leaves half the standard eval playbook describing a component I don't have.&lt;/p&gt;

&lt;p&gt;So I finally sat down and measured the thing I actually built. This post is what that took: pulling the answer path into one function, writing ten questions with known answers, four metrics that each catch a different failure, and span attributes that let a backend score answers as they happen. None of it is elaborate. It's the smallest setup that would tell me when an answer is wrong, and the first run turned up a failure mode that had been sitting there the whole time.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the pipeline looks like
&lt;/h2&gt;

&lt;p&gt;Here's the path a question takes:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The browser records the question and Sarvam's &lt;code&gt;saaras:v3&lt;/code&gt; transcribes it.&lt;/li&gt;
&lt;li&gt;The server loads the paper text and the passage currently being read.&lt;/li&gt;
&lt;li&gt;The model (&lt;code&gt;sarvam-105b&lt;/code&gt; or Claude, whichever I've picked, on my own key) gets the whole paper, the current passage, recent turns, and the question.&lt;/li&gt;
&lt;li&gt;The answer is stripped of anything a voice would read literally, and &lt;code&gt;bulbul:v3&lt;/code&gt; speaks it.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx9ox6vjfju7yj12uwu53.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx9ox6vjfju7yj12uwu53.png" alt="Top row: Spoken question with a microphone icon, then Transcribe, then Paper plus current passage with a document icon, then Model. From Model two curved arrows drop to the bottom row, one to Strip for speech and then Speak answer with a speaker icon, the other to a shaded Trace box with a pen icon and then Score with a tick." width="800" height="447"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The question path across the top, the answer path and the trace below it.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Steps 2 and 3 are the RAG part, and the answer that comes out of them is what this post evaluates.&lt;/p&gt;
&lt;h2&gt;
  
  
  The retrieval step that isn't a vector search
&lt;/h2&gt;

&lt;p&gt;Talkit has no chunking and no embeddings on the answer path, and that was a deliberate choice.&lt;/p&gt;

&lt;p&gt;Uploads are capped at 60,000 characters. The Attention paper, after the reference list is cut, comes to 30,138. That fits in a model's context with plenty of room, so the model sees all of it.&lt;/p&gt;

&lt;p&gt;The bigger reason is the kind of question I actually ask it. Halfway through a paper I'll say "why does that matter" or "what's this number". Those questions retrieve badly by similarity, because the words that matter ("that", "this") point at what was just read out, not at anything an embedding can match. So the current passage goes in as its own labelled block. That passage &lt;em&gt;is&lt;/em&gt; the retrieval signal.&lt;/p&gt;

&lt;p&gt;Every question sends the whole paper, so input tokens scale with paper length, not question complexity. The 60k cap also rules out books and long theses. If Talkit ever answers questions across a whole library, or takes documents past the cap, that's when I'd go add a vector store, and neither of those is true today.&lt;/p&gt;
&lt;h2&gt;
  
  
  Pulling the pipeline out of the route
&lt;/h2&gt;

&lt;p&gt;The answer logic used to live inline in the &lt;code&gt;/api/ask&lt;/code&gt; handler. An eval can't call a FastAPI route without a database, a session, and a stored key. And an eval that &lt;em&gt;reimplements&lt;/em&gt; the prompt assembly measures a copy, which drifts the first time I edit the original.&lt;/p&gt;

&lt;p&gt;So it moved into &lt;code&gt;app/qa.py&lt;/code&gt;, and the route calls it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;answer_question&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;provider&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;TextProvider&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;lang&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;paper&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                          &lt;span class="n"&gt;passage&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                          &lt;span class="n"&gt;history&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                          &lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;tracer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;start_as_current_span&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ask.answer&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;span&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;messages&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;build_messages&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;paper&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;passage&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;history&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="p"&gt;[],&lt;/span&gt; &lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;raw&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;provider&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;converse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;answer_system&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;lang&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                                      &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;ANSWER_MAX_TOKENS&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;answer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;speakable&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;raw&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="nf"&gt;describe_llm_span&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;span&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ask&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;provider&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;provider&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;provider&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;label&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="n"&gt;answer&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;grounding&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;passage&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;paper&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;metadata&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;lang&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;lang&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;history_turns&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;history&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="p"&gt;[])},&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;answer&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;build_messages&lt;/code&gt; puts the paper and passage first, then up to six recent turns, then the question:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;build_messages&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;paper&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;passage&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;history&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
                   &lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
         &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;FULL PAPER:&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;paper&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="s"&gt;CURRENT PASSAGE:&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;passage&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;assistant&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
         &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Understood. I have the paper and I know where the listener is.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;turn&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;history&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;HISTORY_TURNS&lt;/span&gt;&lt;span class="p"&gt;:]:&lt;/span&gt;
        &lt;span class="n"&gt;role&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;turn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;assistant&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;role&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                         &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;turn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;""&lt;/span&gt;&lt;span class="p"&gt;))[:&lt;/span&gt;&lt;span class="n"&gt;HISTORY_TURN_CHARS&lt;/span&gt;&lt;span class="p"&gt;]})&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;()})&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;messages&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And &lt;code&gt;speakable&lt;/code&gt; removes markdown plus a restated question. The system prompt already forbids markdown and the model mostly complies, which isn't good enough for text that goes straight into a text-to-speech voice.&lt;/p&gt;

&lt;p&gt;The refactor was pure motion: behavior didn't change, and the existing 85 tests still pass. I added six more for this module, which brings the offline suite to 91.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pointing OpenTelemetry traces at Confident AI
&lt;/h2&gt;

&lt;p&gt;Talkit already sends OpenTelemetry traces. Spans are created unconditionally, and with no exporter configured the global provider is a no-op. Picking a backend is environment variables, with no vendor SDK in the code.&lt;/p&gt;

&lt;p&gt;I picked &lt;a href="https://www.confident-ai.com/docs" rel="noopener noreferrer"&gt;Confident AI&lt;/a&gt; because &lt;a href="https://deepeval.com/docs/getting-started" rel="noopener noreferrer"&gt;DeepEval&lt;/a&gt;, which I was already using for the evals, is theirs, so the eval runs and the traces land in the same project. &lt;a href="https://langfuse.com/docs/opentelemetry/get-started" rel="noopener noreferrer"&gt;Langfuse&lt;/a&gt;, &lt;a href="https://github.com/Arize-ai/phoenix" rel="noopener noreferrer"&gt;Phoenix&lt;/a&gt; and others accept OpenTelemetry too, and everything up to this point would work with any of them. Pointing it at Confident AI:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;OTEL_EXPORTER_OTLP_ENDPOINT&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;https://otel.confident-ai.com
&lt;span class="nv"&gt;OTEL_EXPORTER_OTLP_HEADERS&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;x-confident-api-key&lt;span class="o"&gt;=&lt;/span&gt;&amp;lt;key&amp;gt;
&lt;span class="nv"&gt;OTEL_LOGS_EXPORTER&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;none
&lt;span class="nv"&gt;OTEL_RESOURCE_ATTRIBUTES&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;confident.trace.environment&lt;span class="o"&gt;=&lt;/span&gt;production
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The HTTP exporter appends &lt;code&gt;/v1/traces&lt;/code&gt; to the endpoint, so the base URL is enough. The environment label rides on the standard resource attribute variable, which the SDK already reads.&lt;/p&gt;

&lt;p&gt;The third line needed a code change. Talkit ships logs through the same exporter so each log record carries its trace and span id. &lt;a href="https://www.confident-ai.com/docs/integrations/opentelemetry" rel="noopener noreferrer"&gt;Confident AI's OpenTelemetry docs&lt;/a&gt; say the endpoint takes traces and not logs, so the logs have to be turned off separately.  &lt;code&gt;OTEL_LOGS_EXPORTER&lt;/code&gt; is the standard variable for this, so the setup now respects it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;ENDPOINT&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;OTEL_LOGS_EXPORTER&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;none&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;logs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;LoggerProvider&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;resource&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;resource&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;logs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add_log_record_processor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;BatchLogRecordProcessor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;OTLPLogExporter&lt;/span&gt;&lt;span class="p"&gt;()))&lt;/span&gt;
    &lt;span class="nf"&gt;set_logger_provider&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;logs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;logging&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getLogger&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;addHandler&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;LoggingHandler&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;logger_provider&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;logs&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;At this point traces arrive, but they're only timings. A trace backend can show that &lt;code&gt;ask.answer&lt;/code&gt; took four seconds. It can't tell whether the answer was grounded in the paper.&lt;/p&gt;

&lt;h2&gt;
  
  
  Making a span scoreable
&lt;/h2&gt;

&lt;p&gt;To score an answer, Confident AI needs to know which span is the model call and what went in and came out. It reads that from span attributes in its own namespace. Talkit writes them in one function in &lt;code&gt;app/telemetry.py&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;describe_llm_span&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;span&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;provider&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                      &lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;grounding&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
                      &lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;metadata&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                      &lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;bool&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                      &lt;span class="n"&gt;metric_collection&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;content&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;TRACE_CONTENT&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;content&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="n"&gt;content&lt;/span&gt;
    &lt;span class="n"&gt;collection&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;TRACE_METRIC_COLLECTION&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;metric_collection&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="n"&gt;metric_collection&lt;/span&gt;

    &lt;span class="n"&gt;span&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set_attributes&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;confident.span.type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;llm&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;confident.trace.name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gen_ai.request.model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;unknown&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gen_ai.provider.name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;provider&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
    &lt;span class="p"&gt;})&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;span&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set_attribute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;confident.trace.user_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;metadata&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;span&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set_attribute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;confident.span.metadata&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;metadata&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt;
    &lt;span class="n"&gt;span&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set_attributes&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;confident.span.input&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;confident.span.output&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;confident.span.retrieval_context&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;grounding&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;confident.trace.input&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;confident.trace.output&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;})&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;collection&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;span&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set_attribute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;confident.span.metric_collection&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;collection&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A few things here cost me more time than the code suggests.&lt;/p&gt;

&lt;p&gt;Text values are JSON-encoded. The convention expects &lt;code&gt;"\"Why scale it?\""&lt;/code&gt;, not &lt;code&gt;"Why scale it?"&lt;/code&gt;, and a list of passages becomes a JSON array string. Before writing this I ran Confident AI's own &lt;code&gt;confident-trace&lt;/code&gt; package against an in-memory exporter and read what it set, so the shapes match what their SDK emits. The collection name is the exception: it's a plain string. A backend accepts a wrongly encoded attribute without complaint and just shows an empty panel, so the tests assert the encoding directly.&lt;/p&gt;

&lt;p&gt;The grounding is the whole paper. &lt;code&gt;retrieval_context&lt;/code&gt; is &lt;code&gt;[passage, paper]&lt;/code&gt;, because that's what the model saw. If I sent only the passage, the faithfulness metric would flag every true claim drawn from elsewhere in the paper as unsupported.&lt;/p&gt;

&lt;p&gt;Content is off unless you turn it on. Scoring needs the question, the answer, and the paper attached to the span and exported. That's my reading and my questions leaving for a third party. So &lt;code&gt;TRACE_CONTENT=true&lt;/code&gt; is an explicit switch. Without it, the span carries the span type, model name, and metadata, and nothing a person wrote.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why attributes and not the SDK
&lt;/h3&gt;

&lt;p&gt;There is an SDK for this, &lt;code&gt;confident-trace&lt;/code&gt;, built on OpenTelemetry. It auto-instruments frameworks like LangChain, LlamaIndex and Pydantic AI, which would be the argument for using it.&lt;/p&gt;

&lt;p&gt;Talkit calls Sarvam over &lt;code&gt;httpx&lt;/code&gt; and Claude through the Anthropic SDK directly, so there is nothing for it to auto-instrument. It would add a dependency and I would still set the input, output and grounding by hand. Writing the attributes myself keeps "no vendor SDK" true, and the ten &lt;code&gt;confident.*&lt;/code&gt; names above are the entire vendor-specific surface: swap backends and the spans, timings and &lt;code&gt;gen_ai.*&lt;/code&gt; attributes go with you. The cost is that a vendor's attribute names now sit in the code, and if they change, that function changes with them.&lt;/p&gt;

&lt;h2&gt;
  
  
  The golden dataset: ten questions about one paper
&lt;/h2&gt;

&lt;p&gt;The live suites in Talkit already use arXiv 1706.03762v7, &lt;em&gt;Attention Is All You Need&lt;/em&gt;, so the evals do too. The goldens live in &lt;code&gt;evals/goldens.json&lt;/code&gt;. Each one has a question, a phrase that locates the current passage, a language, and an expected answer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"question"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Why do they divide by the square root of d k here?"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"passage_contains"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"particular attention"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"lang"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"en"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"expected_output"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"For large values of d k the dot products grow large in magnitude, which pushes the softmax into regions with extremely small gradients. Scaling by one over the square root of d k counteracts that."&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That one does double duty. It tests the decision to send the whole paper: the current passage is the one that &lt;em&gt;introduces&lt;/em&gt; scaled dot-product attention, while the reason for the scaling comes later on, so an answer that only used that passage would miss it.&lt;/p&gt;

&lt;p&gt;The passage is located by searching the chunks for that phrase:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;locate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;needle&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;flat&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;lambda&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sub&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;\s+&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;chunk&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;needle&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;flat&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;chunk&lt;/span&gt;
    &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;SystemExit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;no passage contains &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;needle&lt;/span&gt;&lt;span class="si"&gt;!r}&lt;/span&gt;&lt;span class="s"&gt;; did chunking change?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;An index would silently point at a different passage the first time chunking changes. A missing phrase fails loudly. I checked every phrase against the extracted text before writing the expected answers, and each one matches a passage among the paper's 46.&lt;/p&gt;

&lt;p&gt;What the ten cover:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Six factual questions&lt;/strong&gt; with answers stated in the paper: layers, heads, the positional encoding choice, training time, label smoothing, BLEU.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Two refusals.&lt;/strong&gt; "How much GPU memory did the big model need?" isn't in the paper, and neither is the weather where the authors work. The system prompt says to say so in one sentence rather than fill the gap from general knowledge. These are the only way to check that it does.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One Hinglish question&lt;/strong&gt;, "Encoder mein kitne layers hain?", because Hinglish is a real language option in Talkit with its own prompt, and it can break independently of English.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The scaling question&lt;/strong&gt; above.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Ten questions doesn't cover much of anything. I'll add a golden every time a real answer comes back wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  Four metrics, each for a different failure
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;metrics_for&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="nc"&gt;FaithfulnessMetric&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;threshold&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.8&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="nc"&gt;AnswerRelevancyMetric&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;threshold&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.7&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="nc"&gt;GEval&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Correct&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;threshold&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;evaluation_params&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;SingleTurnParams&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;INPUT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                               &lt;span class="n"&gt;SingleTurnParams&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ACTUAL_OUTPUT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                               &lt;span class="n"&gt;SingleTurnParams&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;EXPECTED_OUTPUT&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
            &lt;span class="n"&gt;evaluation_steps&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Check that every fact in the expected output appears in the actual output. Wording may differ.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Penalise any number, name or claim in the actual output that contradicts the expected output.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;If the expected output says the paper does not address the question, the actual output must say so and must not supply an answer from general knowledge.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;The actual output must be in the same language and script as the expected output.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="nc"&gt;GEval&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Speakable&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;threshold&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;evaluation_params&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;SingleTurnParams&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ACTUAL_OUTPUT&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
            &lt;span class="n"&gt;evaluation_steps&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;The output will be read aloud by a text-to-speech voice to someone with no screen.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;It should be two to four sentences of plain prose.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Penalise lists, headings, markdown, LaTeX, and symbols a voice would read literally, such as square root signs, subscripts or underscores.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Numbers and variable names should be written the way a person would say them.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Faithfulness&lt;/strong&gt; extracts the claims in the answer and checks each against the grounding, which is how it catches invented numbers.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Answer relevancy&lt;/strong&gt; catches an accurate answer to a different question, a real risk when the question is "what does this mean" and "this" could be several things.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Correct&lt;/strong&gt; compares against the expected answer, and it's the metric that scores the refusals. Faithfulness can't: "the paper doesn't say" makes no claims to check.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Speakable&lt;/strong&gt; is the one no generic metric covers. A Talkit answer is heard, not read. A correct answer containing &lt;code&gt;√dk&lt;/code&gt; hands the text-to-speech voice a symbol where it needs words, which is why the prompt asks for spoken forms and this metric checks that it got them.&lt;/p&gt;

&lt;p&gt;Anything code can check, code checks first, and prints before any judge weighs in:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;mechanical_failures&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;lang&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="n"&gt;failures&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
    &lt;span class="n"&gt;sentences&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;(?&amp;lt;=[.!?])\s+&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;()]&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sentences&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;failures&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sentences&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; sentences&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;lang&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;hinglish&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;DEVANAGARI&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;search&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;failures&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Devanagari in a Hinglish answer&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;search&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;[√∈_\\$]&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;failures&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;unspeakable symbol&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;failures&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Hinglish is romanised Hindi, and both Hinglish prompts forbid Devanagari outright because a model asked for it can drift into Hindi script, which a regex catches without any judge weighing in.&lt;/p&gt;

&lt;h2&gt;
  
  
  A judge from a different family
&lt;/h2&gt;

&lt;p&gt;Every metric above is a model grading a model. &lt;a href="https://arxiv.org/abs/2306.05685" rel="noopener noreferrer"&gt;Zheng et al. (2023)&lt;/a&gt; documented that LLM judges show self-enhancement bias: they favor answers written by the same model. So the judge defaults to a different family from the one answering. Talkit's default text model is &lt;code&gt;sarvam-105b&lt;/code&gt;, and the default judge is &lt;code&gt;claude-sonnet-5&lt;/code&gt;. Evaluating the Claude path instead trips a warning that suggests an OpenAI judge:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;provider&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;judge&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;startswith&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;warning: Claude is judging Claude. Pass --judge gpt-5.4 for a &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
          &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;judge from a different family.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The Claude judge broke on its first run. DeepEval's Anthropic wrapper defaults to 1,024 output tokens. Faithfulness starts by listing every factual statement in the grounding, and the grounding here is a whole paper. That list ran past 1,024 tokens, came back as truncated JSON, and DeepEval reported "Evaluation LLM outputted an invalid JSON. Please use a better evaluation model." The fix was the budget:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;judge&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;startswith&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nc"&gt;AnthropicModel&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;generation_kwargs&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;max_tokens&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;16000&lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nc"&gt;OpenAIModel&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Next time a judge hands back broken JSON over long context I'll check the token limit before I blame the model.&lt;/p&gt;

&lt;h2&gt;
  
  
  Running it
&lt;/h2&gt;

&lt;p&gt;Four lines, run by hand when I've changed something worth checking:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt; evals/requirements.txt
curl &lt;span class="nt"&gt;-L&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; attention.pdf https://arxiv.org/pdf/1706.03762v7
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;SARVAM_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;... &lt;span class="nv"&gt;OPENAI_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;...
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;CONFIDENT_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;...
python evals/run.py &lt;span class="nt"&gt;--pdf&lt;/span&gt; attention.pdf &lt;span class="nt"&gt;--judge&lt;/span&gt; gpt-5.4
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The script loads the paper exactly as an upload does (extract, cut the back matter, cap at 60k), answers all ten goldens through &lt;code&gt;answer_question&lt;/code&gt;, runs the mechanical checks, then hands the test cases to DeepEval. With &lt;code&gt;CONFIDENT_API_KEY&lt;/code&gt; set, the run uploads to Confident AI tagged with the provider, the model, the judge, and a short hash of the answer prompt. That hash is what makes prompt changes comparable, since two runs with different hashes can be put side by side against the same goldens.&lt;/p&gt;

&lt;p&gt;Evals have their own requirements file, which includes the app's requirements plus &lt;code&gt;deepeval==4.2.3&lt;/code&gt;. The Docker image never installs DeepEval. One side effect: DeepEval needs &lt;code&gt;python-dotenv&lt;/code&gt; 1.1.1 or later, so the app's pin moved up from 1.0.1.&lt;/p&gt;

&lt;p&gt;Passages in the eval come from the extracted text. The app narrates a cleaned read-aloud script instead, and cleaning a paper takes minutes of paid API calls, which isn't what this suite measures. Wording differs slightly from what I actually hear, and I've accepted that.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the first run found
&lt;/h2&gt;

&lt;p&gt;I ran it with &lt;code&gt;sarvam-105b&lt;/code&gt; answering and &lt;code&gt;gpt-5.4&lt;/code&gt; judging:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Average&lt;/th&gt;
&lt;th&gt;Passed&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Faithfulness&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;9 of 9&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Answer relevancy&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;td&gt;9 of 9&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Correct&lt;/td&gt;
&lt;td&gt;0.91&lt;/td&gt;
&lt;td&gt;8 of 9&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Speakable&lt;/td&gt;
&lt;td&gt;0.93&lt;/td&gt;
&lt;td&gt;9 of 9&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Nine of nine, because one question never got an answer.&lt;/p&gt;

&lt;p&gt;Sometimes there's no answer at all. &lt;code&gt;sarvam-105b&lt;/code&gt; is a reasoning model, and on one question it spent the entire 3,000-token budget reasoning and never wrote a word. The provider raises &lt;code&gt;ChatBudgetExceeded&lt;/code&gt; with &lt;code&gt;finish_reason=length&lt;/code&gt;. I ran the suite four times. Three runs had exactly one unanswered question (the BLEU question twice, the attention-heads question once) and one run answered all ten. The failure is intermittent and doesn't track any particular question. The same thing happened through the app itself, where "How many layers are in the encoder?" came back as a 502. Before evals, the only live test asked one question, so whether it passed depended on the run. The first version of the eval script crashed on it and threw away nine paid answers, so it now records an unanswered question as a failed case with the reason attached.&lt;/p&gt;

&lt;p&gt;Faithfulness didn't discriminate. Every answer scored 1.00. The answers stayed inside the paper, including the refusals. Asked how much GPU memory the big model needed, it said the paper doesn't say, and added that the model has 213 million parameters, which &lt;em&gt;is&lt;/em&gt; in Table 3. That's a good result, but it means faithfulness alone would have told me nothing, and the metric that actually moved was Correct.&lt;/p&gt;

&lt;p&gt;The one Correct failure was my golden's fault. The Hinglish question asks "Encoder mein kitne layers hain?", how many layers are in the encoder. The answer said six. My expected output also listed the two sub-layers, which the question never asked for, and the judge docked it to 0.65 for leaving them out. The judge's reason spelled it out. The golden needs fixing, and the prompt can stay as it is.&lt;/p&gt;

&lt;p&gt;Speakable let one thing through that I'd call a miss: the answer to the scaling question writes "dk" for d k, which a voice reads as a two-letter word. The judge gave it 0.83, comfortably above threshold, and the mechanical check doesn't look for it either. Later answers to the same question through the app did it again. That's the next check to add in code, since a regex settles it better than a judge.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fypdvewi6zh1k54to4eyz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fypdvewi6zh1k54to4eyz.png" alt="Confident AI test run overview named ask-sarvam-ce30d9e6. A dial reads 89 percent, 8 of 9 test cases passing. Metric tiles show Correct 0.91 with one failure, Speakable 0.93, Faithfulness 1.00 and Answer Relevancy 1.00. A hyperparameters panel lists answer_prompt ce30d9e6, judge gpt-5.4, model sarvam-105b, paper_chars 30138 and provider sarvam." width="799" height="383"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;One eval run: the nine questions that got answered, four metrics each, and the hyperparameters that produced them.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcuufxc4it5dikq5ryxxu.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcuufxc4it5dikq5ryxxu.png" alt="Confident AI test case detail for the Hinglish question, Encoder mein kitne layers hain. Correct scored 0.65 against a 0.7 threshold and is marked Failed, while Speakable scored 0.84 and Answer Relevancy 1.00. The judge's written reason says the answer correctly gives six identical layers but omits the two sub-layers the expected output required." width="800" height="413"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The judge's written reason names the detail my expected answer asked for and the question never did.&lt;/em&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  Why it isn't in CI
&lt;/h2&gt;

&lt;p&gt;Talkit's offline checks are pyflakes, 91 tests, and a Node script that checks the voice command matcher against the shipped HTML, and none of them cost anything.&lt;/p&gt;

&lt;p&gt;This suite costs credits on every run: ten answers, then four metrics per answer, and each metric makes several judge calls. And as those runs showed, the model under test sometimes doesn't answer at all. A hard gate would turn that into a flaky build I'd just rerun until it went green.&lt;/p&gt;

&lt;p&gt;So it runs by hand, like the two live suites that exercise Sarvam's speech APIs. The script exits non-zero if any case fails a metric, a mechanical check, or gets no answer, so it could run in a scheduled job later. For now it's a tool I run before changing a prompt or a model.&lt;/p&gt;
&lt;h2&gt;
  
  
  Scoring answers as they happen, with a metric collection
&lt;/h2&gt;

&lt;p&gt;Offline evals only know the questions I thought of. Actually using the thing turns up the ones I didn't.&lt;/p&gt;

&lt;p&gt;Confident AI can score incoming traces with a metric collection: a named set of metrics defined in the project. Talkit reads the name from the environment:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;TRACE_CONTENT&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;true
&lt;/span&gt;&lt;span class="nv"&gt;TRACE_METRIC_COLLECTION&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;talkit-answers
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With both set, every &lt;code&gt;ask.answer&lt;/code&gt; span carries the question, the answer, the grounding, and the collection name, and Confident AI runs the collection's metrics against it. Faithfulness and answer relevancy both work here. Correctness doesn't, because a question asked in the moment has no expected answer written for it. Correctness still comes from goldens. When a live answer scores badly, it becomes a golden.&lt;/p&gt;

&lt;p&gt;One trap when creating the collection: Confident AI offers multi-turn metrics like Turn Faithfulness and Turn Relevancy alongside the single-turn ones. A collection of turn metrics has nothing to score on these spans. Talkit's answer span is a single turn, one question and one answer with its grounding, so the collection needs the single-turn Faithfulness and Answer Relevancy.&lt;/p&gt;

&lt;p&gt;The other thing that cost me an evening: online evaluation needs a judge model of your own. Confident AI includes 100 traces of online evaluation, and after that the trace still arrives, still shows the collection name on the span, and simply is not scored until you configure an evaluation model in the project. The span looked correctly wired for a while before I noticed the banner saying so.&lt;/p&gt;

&lt;p&gt;With a judge configured, the first scored answer came back at 1.00 on both faithfulness and answer relevancy, with a written reason on each, against the same passage and paper the model had been given. The scores in the collection default to a threshold of 0.5, while the offline suite uses 0.8 for faithfulness and 0.7 for relevancy, so the same answer can pass in one place and fail in the other. I haven't reconciled them yet.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftqxnuzg5aew0rkfp9q1x.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftqxnuzg5aew0rkfp9q1x.png" alt="A scored ask.answer span in Confident AI, tagged with the talkit-answers metric collection. Faithfulness and Answer Relevancy both scored 1.00, judged by gpt-5.4." width="799" height="413"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;One answer from the running app, scored on arrival against the passage and paper the model was given.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F91sp05fv0x7v0slkpn0p.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F91sp05fv0x7v0slkpn0p.png" alt="The same trace's Evaluations panel. Answer Relevancy and Faithfulness both read 1.00 against a 0.5 threshold, each with a written reason. Expected Output is None and Retrieval Context is a list of strings holding the paper text." width="799" height="411"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Expected Output is empty, which is the whole reason correctness can't be scored here. Note the 0.5 thresholds against the suite's 0.8 and 0.7.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;When one does score badly, the span has what I need to start: the model and provider, the language, how many history turns were in play, and the exact grounding. The usual first question is whether the problem is in the current passage or the answer. Unlike a vector-search RAG system, there's no "wrong chunks retrieved" branch to rule out, since the model had the whole paper.&lt;/p&gt;

&lt;h2&gt;
  
  
  What else turned up
&lt;/h2&gt;

&lt;p&gt;Running the whole thing locally, upload to answer, surfaced two problems outside the answer path.&lt;/p&gt;

&lt;p&gt;Cleaning falls back more than I knew. Uploading the Attention paper through Sarvam took 837 seconds, and 14 of its 28 cleaning batches fell back to raw text. The log shows the same cause 52 times: a batch overran the token budget and was halved, until it hit the split limit and kept the raw text. That's the same reasoning-budget problem as the missing answers, on the cleaning side. A fallback is recorded and never surfaced, so I just hear citation markers read out with no explanation. It was already on Talkit's list of known gaps. Now it has a number.&lt;/p&gt;

&lt;p&gt;The title is wrong. Talkit's title extractor reads the Attention paper's title as "Provided proper attribution is provided, Google hereby grants permission to", which is the permission banner at the top of the arXiv version.&lt;/p&gt;

&lt;p&gt;Neither is fixed. I only ran into them because I was reading logs and traces for a different reason.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>rag</category>
      <category>llm</category>
      <category>voiceai</category>
    </item>
    <item>
      <title>See O2 AI Assistant in action!</title>
      <dc:creator>Gorakhnath Yadav</dc:creator>
      <pubDate>Thu, 14 May 2026 10:27:46 +0000</pubDate>
      <link>https://dev.to/gorakh13/see-o2-ai-assistant-in-action-4hl6</link>
      <guid>https://dev.to/gorakh13/see-o2-ai-assistant-in-action-4hl6</guid>
      <description>&lt;div class="crayons-card c-embed text-styles text-styles--secondary"&gt;
    &lt;div class="c-embed__content"&gt;
        &lt;div class="c-embed__cover"&gt;
          &lt;a href="https://dev.to/openobserve/i-built-a-dashboard-in-30-seconds-with-ai-500p" class="c-link align-middle" rel="noopener noreferrer"&gt;
            &lt;img alt="" src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fimg.youtube.com%2Fvi%2Ftsm2aDoDxv8%2Fmaxresdefault.jpg" height="auto" class="m-0"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="c-embed__body"&gt;
        &lt;h2 class="fs-xl lh-tight"&gt;
          &lt;a href="https://dev.to/openobserve/i-built-a-dashboard-in-30-seconds-with-ai-500p" rel="noopener noreferrer" class="c-link"&gt;
            I Built a Dashboard in 30 Seconds with AI - DEV Community
          &lt;/a&gt;
        &lt;/h2&gt;
          &lt;p class="truncate-at-3"&gt;
            How the OpenObserve AI Assistant builds production-ready dashboards, creates alerts, and finds root cause — all from plain English. Plus, connect it to your IDE via MCP.
          &lt;/p&gt;
        &lt;div class="color-secondary fs-s flex items-center"&gt;
            &lt;img alt="favicon" class="c-embed__favicon m-0 mr-2 radius-0" src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F8j7kvp660rqzt99zui8e.png"&gt;
          dev.to
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
&lt;/div&gt;


</description>
    </item>
    <item>
      <title>[Boost]</title>
      <dc:creator>Gorakhnath Yadav</dc:creator>
      <pubDate>Wed, 29 Apr 2026 15:40:29 +0000</pubDate>
      <link>https://dev.to/gorakh13/-331i</link>
      <guid>https://dev.to/gorakh13/-331i</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/openobserve/openobserve-just-raised-10m-and-launched-observability-30-with-new-ai-capabilities-3ibl" class="crayons-story__hidden-navigation-link"&gt;OpenObserve Just Raised $10M and Launched Observability 3.0 with New AI Capabilities&lt;/a&gt;
    &lt;div class="crayons-article__cover crayons-article__cover__image__feed"&gt;
      &lt;iframe src="https://www.youtube.com/embed/pdcecXf5zbE" title="OpenObserve Just Raised $10M and Launched Observability 3.0 with New AI Capabilities"&gt;&lt;/iframe&gt;
    &lt;/div&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;
          &lt;a class="crayons-logo crayons-logo--l" href="/openobserve"&gt;
            &lt;img alt="OpenObserve logo" src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Forganization%2Fprofile_image%2F13191%2Fa8a5832e-60af-4ee0-a047-091276bd76be.png" class="crayons-logo__image" width="507" height="498"&gt;
          &lt;/a&gt;

          &lt;a href="/huddle_c62a12743" class="crayons-avatar  crayons-avatar--s absolute -right-2 -bottom-2 border-solid border-2 border-base-inverted  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3902477%2Fff8dc9ed-7a95-4df8-a28a-9a48f1b32eb8.png" alt="huddle_c62a12743 profile" class="crayons-avatar__image" width="800" height="800"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/huddle_c62a12743" class="crayons-story__secondary fw-medium m:hidden"&gt;
              Sara
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                Sara
                
              
              &lt;div id="story-author-preview-content-3583040" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/huddle_c62a12743" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3902477%2Fff8dc9ed-7a95-4df8-a28a-9a48f1b32eb8.png" class="crayons-avatar__image" alt="" width="800" height="800"&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;Sara&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

            &lt;span&gt;
              &lt;span class="crayons-story__tertiary fw-normal"&gt; for &lt;/span&gt;&lt;a href="/openobserve" class="crayons-story__secondary fw-medium"&gt;OpenObserve&lt;/a&gt;
            &lt;/span&gt;
          &lt;/div&gt;
          &lt;a href="https://dev.to/openobserve/openobserve-just-raised-10m-and-launched-observability-30-with-new-ai-capabilities-3ibl" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Apr 29&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/openobserve/openobserve-just-raised-10m-and-launched-observability-30-with-new-ai-capabilities-3ibl" id="article-link-3583040"&gt;
          OpenObserve Just Raised $10M and Launched Observability 3.0 with New AI Capabilities
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag crayons-tag--filled  " href="/t/news"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;news&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/devops"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;devops&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/cloud"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;cloud&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/observability"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;observability&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
          &lt;a href="https://dev.to/openobserve/openobserve-just-raised-10m-and-launched-observability-30-with-new-ai-capabilities-3ibl" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left"&gt;
            &lt;div class="multiple_reactions_aggregate"&gt;
              &lt;span class="multiple_reactions_icons_container"&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/sparkle-heart-5f9bee3767e18deb1bb725290cb151c25234768a0e9a2bd39370c382d02920cf.svg" width="24" height="24"&gt;
                  &lt;/span&gt;
              &lt;/span&gt;
              &lt;span class="aggregate_reactions_counter"&gt;3&lt;span class="hidden s:inline"&gt;&amp;nbsp;reactions&lt;/span&gt;&lt;/span&gt;
            &lt;/div&gt;
          &lt;/a&gt;
            &lt;a href="https://dev.to/openobserve/openobserve-just-raised-10m-and-launched-observability-30-with-new-ai-capabilities-3ibl#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              1&lt;span class="hidden s:inline"&gt;&amp;nbsp;comment&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            1 min read
          &lt;/small&gt;
            
              &lt;span class="bm-initial"&gt;
                

              &lt;/span&gt;
              &lt;span class="bm-success"&gt;
                

              &lt;/span&gt;
            
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


</description>
    </item>
    <item>
      <title>Open-Source Payments: Modular, Flexible, and Built for You</title>
      <dc:creator>Gorakhnath Yadav</dc:creator>
      <pubDate>Tue, 11 Mar 2025 06:57:19 +0000</pubDate>
      <link>https://dev.to/hyperswitchio/open-source-payments-modular-flexible-and-built-for-you-12g7</link>
      <guid>https://dev.to/hyperswitchio/open-source-payments-modular-flexible-and-built-for-you-12g7</guid>
      <description>&lt;p&gt;&lt;iframe width="710" height="399" src="https://www.youtube.com/embed/SWAaMmRFshU"&gt;
&lt;/iframe&gt;
&lt;br&gt;
&lt;strong&gt;Imagine a Linux-like foundation for payments.&lt;/strong&gt; A unified and open-source intelligence that transforms payment agility into a competitive advantage.  &lt;/p&gt;

&lt;p&gt;Introducing &lt;strong&gt;Hyperswitch&lt;/strong&gt;, a modular, open-source, and interoperable full-stack payments platform engineered to put control back in your hands.  &lt;/p&gt;

&lt;h3&gt;
  
  
  Why Hyperswitch?
&lt;/h3&gt;

&lt;p&gt;Enterprises can &lt;strong&gt;build, augment, or customize&lt;/strong&gt; their payment stack with modular solutions. Whether it's &lt;strong&gt;enabling alternate payment methods, streamlining operations across multiple providers and acquirers, or optimizing payment routing&lt;/strong&gt;—Hyperswitch provides the flexibility businesses need.  &lt;/p&gt;

&lt;p&gt;With &lt;strong&gt;advanced analytics on processing costs, seamless authentication, token storage, and reconciliation&lt;/strong&gt;, it simplifies and secures global payments.  &lt;/p&gt;

&lt;h3&gt;
  
  
  Backed by a Global Engineering Force
&lt;/h3&gt;

&lt;p&gt;Hyperswitch is powered by &lt;strong&gt;1,000+ payment engineers&lt;/strong&gt; globally, enabling infrastructure that processes &lt;strong&gt;up to 175 million transactions a day&lt;/strong&gt; with &lt;strong&gt;99.999% availability&lt;/strong&gt;.  &lt;/p&gt;

&lt;p&gt;Engineered with robust &lt;strong&gt;system design principles&lt;/strong&gt;, its &lt;strong&gt;open-source architecture&lt;/strong&gt; ensures &lt;strong&gt;transparency, security, and enterprise-grade reliability&lt;/strong&gt; with &lt;strong&gt;PCI compliance, tokenization, and security certifications&lt;/strong&gt;.  &lt;/p&gt;

&lt;h3&gt;
  
  
  Freedom from Closed-Loop Systems
&lt;/h3&gt;

&lt;p&gt;With Hyperswitch, merchants gain &lt;strong&gt;rapid integration capabilities&lt;/strong&gt; and &lt;strong&gt;freedom from vendor lock-ins&lt;/strong&gt;. The &lt;strong&gt;open-source model&lt;/strong&gt; allows for &lt;strong&gt;risk-free due diligence&lt;/strong&gt;, eliminating the need for lengthy and costly &lt;strong&gt;RFP cycles&lt;/strong&gt;.  &lt;/p&gt;

&lt;p&gt;Merchants can &lt;strong&gt;deploy Hyperswitch in local or cloud environments&lt;/strong&gt; or leverage &lt;strong&gt;Juspay’s hosted sandboxes&lt;/strong&gt; for seamless experimentation.  &lt;/p&gt;

&lt;h3&gt;
  
  
  Take Control of your Payments
&lt;/h3&gt;

&lt;p&gt;Whether you aim to &lt;strong&gt;optimize costs, expand into new markets, or reduce vendor reliance&lt;/strong&gt;, Hyperswitch empowers you to &lt;strong&gt;build what you want, how you want&lt;/strong&gt;.  &lt;/p&gt;

&lt;p&gt;🔗 &lt;strong&gt;&lt;a href="https://github.com/juspay/hyperswitch" rel="noopener noreferrer"&gt;Learn more about open-source payments and Juspay Hyperswitch.&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>opensource</category>
      <category>rust</category>
      <category>community</category>
    </item>
    <item>
      <title>Transitioning from Kubernetes to EC2 for Enhanced Kafka Performance</title>
      <dc:creator>Gorakhnath Yadav</dc:creator>
      <pubDate>Thu, 06 Feb 2025 14:07:56 +0000</pubDate>
      <link>https://dev.to/hyperswitchio/transitioning-from-kubernetes-to-ec2-for-enhanced-kafka-performance-2nmi</link>
      <guid>https://dev.to/hyperswitchio/transitioning-from-kubernetes-to-ec2-for-enhanced-kafka-performance-2nmi</guid>
      <description>&lt;p&gt;Scaling distributed systems is never just about performance, it's also about cost and operational efficiency. While Kubernetes provided us with a solid foundation for container orchestration, we hit unexpected roadblocks when running Kafka clusters at scale.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Costs were rising due to inefficiencies in resource allocation.&lt;/li&gt;
&lt;li&gt;Auto-scaling wasn’t handling our stateful workload well.&lt;/li&gt;
&lt;li&gt;Kafka node management with Strimzi led to operational complexity.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;After months of firefighting, we decided to move from Kubernetes to EC2, a transition that improved performance, simplified operations, and cut costs by 28%.&lt;/p&gt;

&lt;p&gt;Here’s the story of that journey, what worked, and the lessons we learned.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Kubernetes wasn’t working for us?
&lt;/h2&gt;

&lt;h4&gt;
  
  
  1. Resource Allocation Inefficiencies
&lt;/h4&gt;

&lt;p&gt;Kubernetes dynamically manages resources, but in our case, it led to hidden inefficiencies.&lt;/p&gt;

&lt;p&gt;For example, when allocating 2 CPU cores and 8GB RAM, we observed that the actual provisioned resources were often slightly lower (1.8 CPU cores, 7.5GB RAM). This discrepancy may seem trivial, but at scale, it resulted in significant wasted resources and unexpected cost overruns.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Imagine paying for a full tank of fuel, but your car only gets 90% of it. Over time, those missing liters add up.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fy0n9o469nx7pesiknmh4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fy0n9o469nx7pesiknmh4.png" width="800" height="193"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  2. Auto-Scaling Challenges for Stateless Applications
&lt;/h4&gt;

&lt;p&gt;Kubernetes’ auto-scaling mechanism works well for stateless applications, but Kafka isn’t stateless. When resources ran out, Kubernetes would restart our Kafka application instead of efficiently scaling it.&lt;/p&gt;

&lt;p&gt;This resulted in:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;15-second delays in message processing.&lt;/li&gt;
&lt;li&gt;Increased latency during scaling events.&lt;/li&gt;
&lt;li&gt;Operational headaches managing stateful workloads.&lt;/li&gt;
&lt;/ul&gt;

&lt;h4&gt;
  
  
  3. Kafka Node Management with Strimzi
&lt;/h4&gt;

&lt;p&gt;Initially, we relied on Strimzi for managing Kafka clusters. However, it had major drawbacks:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;New Kafka nodes often failed to integrate properly.&lt;/li&gt;
&lt;li&gt;Manual intervention was required for every scaling event.&lt;/li&gt;
&lt;li&gt;Overall Kafka performance was unpredictable.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;Managing our Kafka clusters felt like playing whack-a-mole every time we solved one issue, another would pop up.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Decided to move to EC2
&lt;/h2&gt;

&lt;p&gt;After evaluating various alternatives, we decided to move Kafka from Kubernetes to EC2. This gave us more control over resource allocation, auto-scaling, and cluster management.&lt;/p&gt;

&lt;p&gt;Here’s what changed:&lt;/p&gt;

&lt;h4&gt;
  
  
  1. Replacing Strimzi with a Custom Kafka Controller
&lt;/h4&gt;

&lt;p&gt;Instead of relying on third-party tools, we built an in-house Kafka Controller tailored to our needs.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Seamless integration of new Kafka nodes&lt;/li&gt;
&lt;li&gt;Automated scaling based on real-time workload analysis&lt;/li&gt;
&lt;li&gt;Better cluster management with minimal manual intervention&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Result&lt;/strong&gt;? Kafka nodes were now automatically recognized and integrated instantly.&lt;/p&gt;

&lt;h4&gt;
  
  
  2. Precise Resource Allocation
&lt;/h4&gt;

&lt;p&gt;Unlike Kubernetes, where we had limited control over resource provisioning, EC2 allowed us to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Allocate exactly the CPU and memory we needed.&lt;/li&gt;
&lt;li&gt;Avoid wasted resources and over-provisioning costs.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Example&lt;/strong&gt;: Previously, we paid $180/month per instance on Kubernetes. After transitioning to EC2, this dropped to $130/month, saving 28% on infrastructure costs.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Frh0mn7w28vbzs2gtnr1r.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Frh0mn7w28vbzs2gtnr1r.png" width="800" height="319"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  3. Streamlined Kafka Node Support
&lt;/h4&gt;

&lt;p&gt;With EC2, we could now:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Scale up Kafka nodes seamlessly without restarts.&lt;/li&gt;
&lt;li&gt;Perform vertical scaling (switching to more powerful machines) with zero downtime.&lt;/li&gt;
&lt;li&gt;Ensure predictable
performance under peak loads.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Last month, we moved from a &lt;strong&gt;T-class instance&lt;/strong&gt; to a &lt;strong&gt;C-class instance&lt;/strong&gt; in &lt;strong&gt;EC2 without downtime&lt;/strong&gt;. If we had been on Kubernetes, this would have required:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Creating a new node group.&lt;/li&gt;
&lt;li&gt;Rebalancing partitions manually.&lt;/li&gt;
&lt;li&gt;Managing potential downtime.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Instead, on EC2, it was a simple instance upgrade with zero complexity, zero downtime.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Impact: Cost Savings &amp;amp; efficiency gains
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fi81sqgpxexjzrpkdbfo2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fi81sqgpxexjzrpkdbfo2.png" width="800" height="165"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Key lessons &amp;amp; Takeaways
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Not all workloads are ideal for Kubernetes – It’s great for general-purpose container orchestration but not always the best for stateful applications like Kafka.&lt;/li&gt;
&lt;li&gt;Custom solutions can be worth it – Building an in-house Kafka Controller gave us better control and reliability.&lt;/li&gt;
&lt;li&gt;Cost inefficiencies add up – Even small inefficiencies in resource allocation can result in thousands of dollars lost at scale.&lt;/li&gt;
&lt;li&gt;EC2 provides better flexibility – We gained granular control over scaling and performance with EC2.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;If you’re running Kafka on Kubernetes and experiencing similar issues, EC2 might be a better fit.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Who should consider moving? Teams struggling with stateful workloads on Kubernetes.&lt;/li&gt;
&lt;li&gt;Who should stick with Kubernetes? Those managing stateless, highly dynamic applications.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Every infrastructure decision should be guided by workload needs, not just industry trends. Kubernetes is powerful, but for Kafka, EC2 provided the right balance of cost, performance, and operational efficiency for us.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Hacktoberfest 2024: Maintainer POV</title>
      <dc:creator>Gorakhnath Yadav</dc:creator>
      <pubDate>Thu, 31 Oct 2024 06:25:09 +0000</pubDate>
      <link>https://dev.to/gorakh13/hacktoberfest-2024-maintainer-pov-4c1d</link>
      <guid>https://dev.to/gorakh13/hacktoberfest-2024-maintainer-pov-4c1d</guid>
      <description>&lt;p&gt;As October comes to an end, I’ve been reflecting on my experience during &lt;a href="https://hyperswitch.io/hacktoberfest" rel="noopener noreferrer"&gt;Hacktoberfest this year&lt;/a&gt;. This time, I participated as a maintainer for our project at &lt;a class="mentioned-user" href="https://dev.to/hyperswitch"&gt;@hyperswitch&lt;/a&gt; , which deepened my appreciation for the importance of community in open source.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;My Role:&lt;/strong&gt; As a DevRel, I was looking after issue creation, guiding people to resources, managing the assignment of issues, and coordinating among the contributors and the internal team for queries and reviews.&lt;/p&gt;

&lt;h3&gt;
  
  
  How It Went:📈
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;We kicked off with a good start; as we created issues in the last weeks of September, contributors were already discussing and working on them.&lt;/li&gt;
&lt;li&gt;As soon as we started, we had a lot of people coming in and asking for assignments of issues so they could work on it. By October 3rd, we had 48 issues assigned, 18 PRs raised.&lt;/li&gt;
&lt;li&gt;Tried to make issues beginner-friendly, given the complexity of digital payments.&lt;/li&gt;
&lt;li&gt;Saw participation from all skill levels, from first-time contributors to seasoned Hacktoberfest participants.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Learnings for us:🤔
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Hacktoberfest allowed me to view our project from a fresh perspective, as contributors noticed details I might overlook.&lt;/li&gt;
&lt;li&gt;Questions from contributors helped identify areas in our documentation needing clarification, which we’ll be updating soon.&lt;/li&gt;
&lt;li&gt;Also updating guidelines to address common queries raised by contributors.&lt;/li&gt;
&lt;li&gt;Participation seemed lower this year, and I wonder if this trend affected other projects as well. I'd love to hear your thoughts on this.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Community is WOW!❤️
&lt;/h3&gt;

&lt;p&gt;Hacktoberfest also highlighted the diversity of the open-source community. As we had people from each level of skills and expertise, each contributor brought unique insights. Some were experienced developers who tackled issues with ease, while others were just starting out but were eager to learn and help. Working with such a diverse group reminded me that open source thrives on this variety of skills and perspectives.&lt;/p&gt;

&lt;p&gt;Being a maintainer also reinforced the responsibility that comes with this role. Open source is not just about code—it’s about creating a space where everyone feels welcome to contribute. It’s essential to be approachable, give helpful feedback, and appreciate the time each person spends on their contribution.&lt;/p&gt;

&lt;h3&gt;
  
  
  Road Ahead:🛣️
&lt;/h3&gt;

&lt;p&gt;This Hacktoberfest was full of learning, connection, and growth for our community. My experience as a maintainer at Hyperswitch provided insights I hadn’t expected and reminded me why open source is about much more than just coding. I’m excited to continue supporting this community and watching it grow.&lt;/p&gt;

&lt;p&gt;Here’s to Hacktoberfest 2024🥂 and to all the contributors who bring life to open source!🚀&lt;/p&gt;

</description>
      <category>hacktoberfest</category>
      <category>opensource</category>
      <category>softwaredevelopment</category>
      <category>softwareengineering</category>
    </item>
    <item>
      <title>🚀 Hyperswitch Plugin Development Hackathon: Build, Compete, and Win Big! 🚀</title>
      <dc:creator>Gorakhnath Yadav</dc:creator>
      <pubDate>Wed, 09 Oct 2024 10:48:02 +0000</pubDate>
      <link>https://dev.to/hyperswitchio/hyperswitch-plugin-development-hackathon-build-compete-and-win-big-nnb</link>
      <guid>https://dev.to/hyperswitchio/hyperswitch-plugin-development-hackathon-build-compete-and-win-big-nnb</guid>
      <description>&lt;p&gt;Hey developers! 🔥&lt;/p&gt;

&lt;p&gt;We’re excited to announce the &lt;a href="https://github.com/juspay/hyperswitch/wiki/Plugin-Development-Hackathon" rel="noopener noreferrer"&gt;Hyperswitch's Plugin Development Hackathon&lt;/a&gt;, running from October 9th to October 31st, 2024. It’s your chance to develop plugins for Hyperswitch, collaborate with fellow developers, and win big rewards!&lt;/p&gt;

&lt;h2&gt;
  
  
  What’s the Hackathon About?
&lt;/h2&gt;

&lt;p&gt;You'll be building plugins for platforms like MedusaJS, CommerceTools, or Prestashop, helping to extend Hyperswitch’s payment orchestration platform. &lt;strong&gt;Participate individually or as a team&lt;/strong&gt; (up to 4 members) and tackle real-world payment challenges while competing for awesome prizes!&lt;/p&gt;

&lt;h3&gt;
  
  
  Key Dates:
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Registration Deadline&lt;/strong&gt;: October 14th, 2024&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Approach Submission&lt;/strong&gt;: October 14th, 2024 (EOD IST)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Final Submission&lt;/strong&gt;: October 31st, 2024&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Demo Day&lt;/strong&gt;: First week of November 2024 (TBA)&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  How to Participate:
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Sign Up&lt;/strong&gt;: Register before October 14th to get started.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Submit Your Approach&lt;/strong&gt;: Outline the core features, integration with Hyperswitch’s API, and edge case handling by October 14th.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Get Shortlisted&lt;/strong&gt;: The Hyperswitch team will review your submission and notify shortlisted teams.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Develop Your Plugin&lt;/strong&gt;: If selected, you’ll have two weeks to complete your plugin.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Demo Day&lt;/strong&gt;: Showcase your plugin in early November, dates to be announced!&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Prizes:
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;MedusaJS Plugin&lt;/strong&gt;: $2000&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CommerceTools Plugin&lt;/strong&gt;: $2000&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Prestashop Plugin&lt;/strong&gt;: $2000&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Ready to showcase your skills? &lt;a href="https://github.com/juspay/hyperswitch/wiki/Plugin-Development-Hackathon" rel="noopener noreferrer"&gt;Register today&lt;/a&gt; and start building your plugin! Let’s code, collaborate, and win together. 💻&lt;/p&gt;

&lt;p&gt;Good luck to all participants! 🎉&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/hyperswitch" rel="noopener noreferrer"&gt;Star Us on GitHub&lt;/a&gt;&lt;br&gt;
&lt;a href="https://x.com/HyperSwitchIO" rel="noopener noreferrer"&gt;Follow us on twitter&lt;/a&gt;&lt;/p&gt;

</description>
      <category>opensource</category>
      <category>hackathon</category>
      <category>fintech</category>
    </item>
    <item>
      <title>Hacktoberfest 2024 with Hyperswitch</title>
      <dc:creator>Gorakhnath Yadav</dc:creator>
      <pubDate>Fri, 27 Sep 2024 09:07:40 +0000</pubDate>
      <link>https://dev.to/hyperswitchio/hacktoberfest-2024-with-hyperswitch-47dk</link>
      <guid>https://dev.to/hyperswitchio/hacktoberfest-2024-with-hyperswitch-47dk</guid>
      <description>&lt;h2&gt;
  
  
  Join the Celebration of Open Source!
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://hyperswitch.io/hacktoberfest" rel="noopener noreferrer"&gt;Hyperswitch&lt;/a&gt; is excited to participate in &lt;strong&gt;Hacktoberfest 2024!&lt;/strong&gt; Join us in building a robust payment infrastructure that serves billions of people worldwide. Contribute to our open-source projects and earn exciting rewards.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to Get Involved
&lt;/h2&gt;

&lt;p&gt;1. Explore Hyperswitch Repository:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Visit our &lt;a href="https://github.com/juspay/hyperswitch" rel="noopener noreferrer"&gt;GitHub page&lt;/a&gt;.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Star and watch the repositories to stay updated with new issues.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;2. Find Issues to Work On:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Look for issues tagged with &lt;a href="https://github.com/juspay/hyperswitch/issues?q=is%3Aopen+is%3Aissue+label%3Ahacktoberfest" rel="noopener noreferrer"&gt;"hacktoberfest"&lt;/a&gt;.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;If you have ideas for improvements, feel free to create new issues,  and inform us using community channels.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;3. Make Your Contribution:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Comment on the issue you want to work on: "I would like to work on this. Can you assign it to me?"&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Ensure you follow up within 3-4 days to show you are actively working on it.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Submit your pull request (PR) and link it to the issue. Tag a maintainer for review.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;4. Get Support:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Our maintainers are here to guide you. Join our Community for support and to connect with fellow contributors.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;5. Earn Swag:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Get one PR merged to earn exclusive Hyperswitch swag.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Get 3 or more PRs merged to receive a special swag kit from Hyperswitch.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fn4rxr0ue5eio3un70bmj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fn4rxr0ue5eio3un70bmj.png" alt="Hyperswitch goodies" width="800" height="480"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Contribute?
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Learning Opportunities: Enhance your skills by working on real-world projects.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Networking: Connect with a global community of developers.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Rewards: Earn cool swag for your contributions.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Join the Hyperswitch Community
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://discord.gg/wJZ7DVW8mm" rel="noopener noreferrer"&gt;Discord&lt;/a&gt; &lt;/li&gt;
&lt;li&gt;
&lt;a href="https://join.slack.com/t/hyperswitch-io/shared_invite/zt-2jqxmpsbm-WXUENx022HjNEy~Ark7Orw" rel="noopener noreferrer"&gt;Slack&lt;/a&gt; &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://github.com/hyperswitch" rel="noopener noreferrer"&gt;Star Us on GitHub&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Happy Contributing!🎉&lt;/p&gt;

</description>
      <category>digitalpayments</category>
      <category>hacktoberfest</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Dive Into the World of Payments - Join Payment Meetup Group’s First Session!</title>
      <dc:creator>Gorakhnath Yadav</dc:creator>
      <pubDate>Mon, 09 Sep 2024 03:29:40 +0000</pubDate>
      <link>https://dev.to/gorakh13/dive-into-the-world-of-payments-join-payment-meetup-groups-first-session-41ja</link>
      <guid>https://dev.to/gorakh13/dive-into-the-world-of-payments-join-payment-meetup-groups-first-session-41ja</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Have you ever wondered what happens behind the scenes when you click "Pay Now"? &lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If you're a developer, payment enthusiast, or just curious about how digital payments via card works, we have an exciting session lined up just for you!&lt;/p&gt;

&lt;p&gt;At &lt;strong&gt;Payments Meetup Group&lt;/strong&gt;, we’re diving deep into the mechanisms that power online transactions via cards, and we want to bring you along for the journey. Our upcoming session, led by Samraat Bansal, Engineering Tech Lead at Hyperswitch, will explore every detail from data encryption and transmission to how various players like issuing banks, card networks, and merchants handle your card payment.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why You Should Join:
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;In-depth Insights: Understand the flow of transactions and how card payments are processed securely.&lt;/li&gt;
&lt;li&gt;Learn from Experts: Samraat Bansal brings hands-on experience from leading Hyperswitch’s tech innovations.&lt;/li&gt;
&lt;li&gt;Engage with Community: This is a perfect opportunity to network with fellow developers and digital payment enthusiasts.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Session Highlights:
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;What happens when you enter your card details and click "Pay Now"?&lt;/li&gt;
&lt;li&gt;The critical role of encryption and transmission in ensuring transaction security.&lt;/li&gt;
&lt;li&gt;How various stakeholders like banks and merchants collaborate to approve or decline your transaction.&lt;/li&gt;
&lt;li&gt;Behind-the-scenes interactions between card networks, acquiring banks, and issuing banks.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  When and how to join?
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Date &amp;amp; Time&lt;/strong&gt;: 11th Sept, 2024, 07:30 PM (GMT +2)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Where&lt;/strong&gt;: Online (&lt;a href="https://www.meetup.com/payment-group/events/303230836/" rel="noopener noreferrer"&gt;Link to register&lt;/a&gt;)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Don’t miss out on this deep dive into the fascinating world of digital payments. Whether you’re a seasoned developer or new to the world of digital finance, this session is tailored to give you valuable insights and practical knowledge.&lt;/p&gt;

&lt;p&gt;RSVP Now and be part of this insightful conversation!&lt;/p&gt;

</description>
      <category>digitalpayments</category>
      <category>fintech</category>
      <category>onlinepayments</category>
    </item>
    <item>
      <title>Securing Digital Transactions: How Hyperswitch makes Payment Protection a Priority</title>
      <dc:creator>Gorakhnath Yadav</dc:creator>
      <pubDate>Thu, 01 Aug 2024 09:26:44 +0000</pubDate>
      <link>https://dev.to/hyperswitchio/securing-digital-transactions-how-hyperswitch-makes-payment-protection-a-priority-3dnb</link>
      <guid>https://dev.to/hyperswitchio/securing-digital-transactions-how-hyperswitch-makes-payment-protection-a-priority-3dnb</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Security is not a product, but a process  - &lt;a href="https://en.wikipedia.org/wiki/Bruce_Schneier" rel="noopener noreferrer"&gt;Bruce Schneier&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;iframe width="710" height="399" src="https://www.youtube.com/embed/UOTJ3-spfBM"&gt;
&lt;/iframe&gt;
&lt;br&gt;
In today's digital landscape, data security is paramount, especially when it comes to online payments. At &lt;a href="https://hyperswitch.io/" rel="noopener noreferrer"&gt;Hyperswitch&lt;/a&gt;, we've built our platform with security as a fundamental principle. Let's explore the robust measures we've implemented to ensure a safe and secure payments infrastructure.&lt;/p&gt;

&lt;h2&gt;
  
  
  Our Top Priority: Customer Security
&lt;/h2&gt;

&lt;p&gt;When a customer makes a payment through &lt;a href="https://hyperswitch.io/" rel="noopener noreferrer"&gt;Hyperswitch&lt;/a&gt;, their card information is immediately encrypted at the source. We use the SSL/TLS 1.2 protocol for transmission, adhering to PCI standards for handling card data.&lt;br&gt;
For customers who opt to store their card details, we've developed a secure Card Vault. This system employs multiple layers of protection:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://blog.gigamon.com/2021/07/14/what-is-tls-1-2-and-why-should-you-still-care/" rel="noopener noreferrer"&gt;SSL/TLS&lt;/a&gt; 1.2 encryption&lt;/li&gt;
&lt;li&gt;JWE + JWS for secure data transmission&lt;/li&gt;
&lt;li&gt;Double encryption for stored data&lt;/li&gt;
&lt;li&gt;A two-key custodian system for enhanced security&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Additionally, all customer Personally Identifiable Information (PII) is TLS encrypted in transit and &lt;a href="https://en.wikipedia.org/wiki/Advanced_Encryption_Standard" rel="noopener noreferrer"&gt;AES&lt;/a&gt; encrypted at rest.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keeping Your Business Secure
&lt;/h2&gt;

&lt;p&gt;For our merchant, we know that not everyone needs to be in on all the secrets. That's why we've built up an access control system that's more discerning. API credentials? We treat those with more care than a rare vintage car, wrapping them in layers of AES encryption. Also, we mask the sensitive merchant details in the dashboard.&lt;/p&gt;

&lt;h2&gt;
  
  
  Hyperswitch's Approach to Security
&lt;/h2&gt;

&lt;p&gt;At &lt;a href="https://hyperswitch.io/" rel="noopener noreferrer"&gt;Hyperswitch&lt;/a&gt;, we're not just building a payment system; we're crafting a digital fortress. Our multi-tenant system comes with unique data encryption keys for each business, creating a security layer that's both impressive and effective.&lt;br&gt;
We've also enlisted Rust as our companion, using its robust type system to keep sensitive information under wraps. It's like having a very diligent, slightly obsessive assistant who never lets a secret slip.&lt;/p&gt;

&lt;h2&gt;
  
  
  Serious Security with a Smile
&lt;/h2&gt;

&lt;p&gt;In the end, we're not just developing an application; we're creating a secure future for digital transactions. Our goal is to allow businesses to focus on their core competencies while we ensure the protection of their customers' financial information.&lt;br&gt;
With &lt;a href="https://hyperswitch.io/" rel="noopener noreferrer"&gt;Hyperswitch&lt;/a&gt;, security isn't just a feature—it's a cornerstone of our service. We're dedicated to maintaining the highest standards of data protection, giving both merchants and customers peace of mind in their digital transactions.&lt;/p&gt;




&lt;p&gt;Want to contribute? Check out some of our &lt;a href="https://github.com/juspay/hyperswitch/issues?q=is%3Aopen+is%3Aissue+label%3A%22good+first+issue%22" rel="noopener noreferrer"&gt;good first issues here.&lt;/a&gt;&lt;br&gt;
Try Hyperswitch. &lt;a href="https://app.hyperswitch.io/" rel="noopener noreferrer"&gt;Get your API keys here.&lt;/a&gt; Happy reading!&lt;/p&gt;

</description>
      <category>hyperswitch</category>
      <category>security</category>
      <category>fintech</category>
    </item>
    <item>
      <title>Scaling to 125 Million Transactions per Day: Juspay's Engineering Principles</title>
      <dc:creator>Gorakhnath Yadav</dc:creator>
      <pubDate>Wed, 19 Jun 2024 14:32:56 +0000</pubDate>
      <link>https://dev.to/hyperswitchio/scaling-to-125-million-transactions-per-day-juspays-engineering-principles-2bj1</link>
      <guid>https://dev.to/hyperswitchio/scaling-to-125-million-transactions-per-day-juspays-engineering-principles-2bj1</guid>
      <description>&lt;p&gt;At Juspay, we process 125 million transactions per day, with peak traffic reaching 5,000 transactions per second, all while maintaining 99.99% uptime. Handling such enormous volumes demands a robust, reliable, and scalable system. In this post, we'll walk you through our core engineering principles and how they've shaped our engineering decisions and systems.&lt;br&gt;
When designing systems at this scale, several challenges naturally arise:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Reliability vs Scale:&lt;/strong&gt; Generally, as you scale, you tend to exhaust resources, which can affect system availability.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reliability vs Agility:&lt;/strong&gt; Frequent releases and system changes can impact system reliability.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Scale vs Cost-Effectiveness:&lt;/strong&gt; Scaling requires more resources, 
leading to higher costs.&lt;/li&gt;
&lt;/ol&gt;
&lt;h3&gt;
  
  
  Core Engineering Pillars
&lt;/h3&gt;

&lt;p&gt;We've been able to strike the right balance between these challenges by anchoring our tech stack on four pillars:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Build zero downtime stacks:&lt;/strong&gt; Solve reliability by building redundancy at each layer to achieve almost 100% uptime.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Horizontally scalable systems:&lt;/strong&gt; Solve scalability by building systems that can scale horizontally by removing bottlenecks.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Build agile systems for frequent bug-free releases.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Build performant systems for low latency, high throughput transaction processing.&lt;/strong&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;h3&gt;
  
  
  Adopting Haskell Programming Language
&lt;/h3&gt;

&lt;p&gt;To achieve our goals, we've made a critical investment: adopting the Haskell programming language. Haskell, a functional programming language, offers performance akin to C, which is closer to the machine and processes transactions much faster. With Haskell, we've reduced our transaction &lt;strong&gt;processing time to less than 100 milliseconds.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Here's an example of a Haskell function that adds two numbers:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;-- Define the add function
add :: Int -&amp;gt; Int -&amp;gt; Int
add x y = x + y

-- Main function to test the add function
main :: IO ()
main = do
    let sum = add 3 5
    putStrLn ("The sum of 3 and 5 is " ++ show sum)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This function showcases Haskell's concise and readable syntax.&lt;/p&gt;

&lt;p&gt;Additionally, Haskell's readability, like English, enables non-technical folks to read the code easily, verify business logic, and sign off on features during development itself. As a strong-typed language, Haskell enforces a set of rules to ensure consistency of results, helping us preempt failures and achieve zero technical declines.&lt;/p&gt;

&lt;h3&gt;
  
  
  Cache-based Shock Absorber
&lt;/h3&gt;

&lt;p&gt;To handle scale and remove database bottlenecks, we introduced a horizontally scalable caching layer where real-time transactions are served from this cache layer and later drained to the database.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Foujubo2bjosui4hnxk7g.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Foujubo2bjosui4hnxk7g.png" alt="Implementation of Redis Shock Absorber" width="800" height="444"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Scaling up and down the cache layer is relatively easy and cost-effective compared to scaling databases.&lt;/p&gt;

&lt;h3&gt;
  
  
  Rapid Deployment and Release Frameworks
&lt;/h3&gt;

&lt;p&gt;With rapid development comes the challenge of frequent production releases. To achieve agility through frequent releases without compromising reliability, we've built internal tools for automated releases with minimal manual effort. These tools monitor the performance of the release by benchmarking error codes against the previous stable version of the codebase:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;ART (Automated Regression Tester):&lt;/strong&gt; A system that records production payloads and runs them in the UAT system against a new deployment to identify bugs early.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Autopilot:&lt;/strong&gt; A tool that creates a new deployment and performs traffic staggering from 1% onwards.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A/B testing 
framework:&lt;/strong&gt; A system that monitors and benchmarks the new deployment's performance against the previous stable version. Based on this benchmark, the system automatically decides to scale up the traffic or abort the deployment.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Hyperswitch: An Open-Source Payments Switch
&lt;/h3&gt;

&lt;p&gt;We're carrying these learnings forward to our latest product, &lt;a href="https://github.com/juspay/hyperswitch" rel="noopener noreferrer"&gt;Hyperswitch&lt;/a&gt;, an open-source payments switch. Every line of code powering our stack is available for you to see. With Hyperswitch, our vision is to ensure every business has access to world-class payment infrastructure.&lt;/p&gt;

&lt;h4&gt;
  
  
  Conclusion
&lt;/h4&gt;

&lt;p&gt;Through these investments, we've built reliable, agile, and scalable systems, enabling our engineers to solve exciting new problems and fostering a culture of systems thinking within the company. We encourage developers to engage with the open-source Hyperswitch project and explore the principles and technologies we've adopted to handle massive scale and high-volume transaction processing.&lt;/p&gt;

</description>
      <category>digitalpayments</category>
      <category>rust</category>
      <category>opensource</category>
      <category>haskell</category>
    </item>
  </channel>
</rss>
