<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: mohamed khaled</title>
    <description>The latest articles on DEV Community by mohamed khaled (@mohamed_khaled_2811).</description>
    <link>https://dev.to/mohamed_khaled_2811</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4174304%2F632910d1-a40e-45a7-bf6c-3c409114f6fb.jpg</url>
      <title>DEV Community: mohamed khaled</title>
      <link>https://dev.to/mohamed_khaled_2811</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/mohamed_khaled_2811"/>
    <language>en</language>
    <item>
      <title>[Boost]</title>
      <dc:creator>mohamed khaled</dc:creator>
      <pubDate>Fri, 09 Oct 2026 22:55:52 +0000</pubDate>
      <link>https://dev.to/mohamed_khaled_2811/-48pb</link>
      <guid>https://dev.to/mohamed_khaled_2811/-48pb</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/mohamed_khaled_2811/stop-guessing-your-rag-hyperparameters-why-i-built-a-local-first-benchmarking-framework-2a95" class="crayons-story__hidden-navigation-link"&gt;Stop Guessing Your RAG Hyperparameters: Why I Built a Local-First Benchmarking Framework&lt;/a&gt;
    &lt;div class="crayons-article__cover crayons-article__cover__image__feed"&gt;
      &lt;iframe src="https://www.youtube.com/embed/SOXkpL4Q9PE" title="Stop Guessing Your RAG Hyperparameters: Why I Built a Local-First Benchmarking Framework"&gt;&lt;/iframe&gt;
    &lt;/div&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/mohamed_khaled_2811" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4174304%2F632910d1-a40e-45a7-bf6c-3c409114f6fb.jpg" alt="mohamed_khaled_2811 profile" class="crayons-avatar__image" width="96" height="96"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/mohamed_khaled_2811" class="crayons-story__secondary fw-medium m:hidden"&gt;
              mohamed khaled
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                mohamed khaled
                
                
              
              &lt;div id="story-author-preview-content-4825728" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/mohamed_khaled_2811" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4174304%2F632910d1-a40e-45a7-bf6c-3c409114f6fb.jpg" class="crayons-avatar__image" alt="" width="96" height="96"&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;mohamed khaled&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/mohamed_khaled_2811/stop-guessing-your-rag-hyperparameters-why-i-built-a-local-first-benchmarking-framework-2a95" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Oct 9&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/mohamed_khaled_2811/stop-guessing-your-rag-hyperparameters-why-i-built-a-local-first-benchmarking-framework-2a95" id="article-link-4825728"&gt;
          Stop Guessing Your RAG Hyperparameters: Why I Built a Local-First Benchmarking Framework
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/ai"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;ai&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/rag"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;rag&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/opensource"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;opensource&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/python"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;python&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
            &lt;a href="https://dev.to/mohamed_khaled_2811/stop-guessing-your-rag-hyperparameters-why-i-built-a-local-first-benchmarking-framework-2a95#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              &lt;span class="hidden s:inline"&gt;Add&amp;nbsp;Comment&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            2 min read
          &lt;/small&gt;
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


</description>
      <category>ai</category>
      <category>llm</category>
      <category>machinelearning</category>
      <category>rag</category>
    </item>
    <item>
      <title>Stop Guessing Your RAG Hyperparameters: Why I Built a Local-First Benchmarking Framework</title>
      <dc:creator>mohamed khaled</dc:creator>
      <pubDate>Fri, 09 Oct 2026 22:54:51 +0000</pubDate>
      <link>https://dev.to/mohamed_khaled_2811/stop-guessing-your-rag-hyperparameters-why-i-built-a-local-first-benchmarking-framework-2a95</link>
      <guid>https://dev.to/mohamed_khaled_2811/stop-guessing-your-rag-hyperparameters-why-i-built-a-local-first-benchmarking-framework-2a95</guid>
      <description>&lt;h1&gt;
  
  
  Stop Guessing Your RAG Hyperparameters: Why I Built a Local-First Benchmarking Framework
&lt;/h1&gt;

&lt;p&gt;Building a "Hello World" Retrieval-Augmented Generation (RAG) app takes about 5 minutes. But taking that pipeline to production and ensuring it consistently gives the right answers? That takes months.&lt;/p&gt;

&lt;p&gt;Every time I tweaked a hyperparameter—changing the retrieval method, adding a reranker, or adjusting the &lt;code&gt;Top-K&lt;/code&gt; value—I found myself playing a guessing game. &lt;em&gt;Did this change actually improve the output, or did it just break a different edge case?&lt;/em&gt; &lt;/p&gt;

&lt;p&gt;I needed a way to measure the impact of these changes systematically, without relying on expensive cloud observability platforms. That’s why I built &lt;strong&gt;&lt;a href="https://github.com/Mohamed28112003/Muffakir" rel="noopener noreferrer"&gt;Muffakir&lt;/a&gt;&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Problem with RAG Development
&lt;/h2&gt;

&lt;p&gt;When evaluating RAG pipelines, developers face three major blind spots:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Lack of Visibility:&lt;/strong&gt; You see the final answer, but you don't easily see the exact context retrieved, the generated queries, or the step-by-step latency.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Experiment Chaos:&lt;/strong&gt; Running 50 trials with different prompt templates and rerankers usually ends up in messy Jupyter notebooks or scattered logs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reproducibility Issues:&lt;/strong&gt; If a configuration works today, can you reproduce the exact same result tomorrow if the underlying documents change?&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Enter Muffakir
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Muffakir&lt;/strong&gt; is an open-source, local-first RAG optimization framework. It acts as your control center for building, benchmarking, and comparing RAG pipelines on your own data. &lt;/p&gt;

&lt;p&gt;Instead of writing custom evaluation scripts for every project, Muffakir provides a unified interface to tune your system and understand exactly &lt;em&gt;why&lt;/em&gt; a specific configuration won.&lt;/p&gt;

&lt;h3&gt;
  
  
  Key Features Under the Hood:
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;ComposerUI:&lt;/strong&gt; A dedicated local dashboard that lets you define your search space. You can swap out retrievers, test different rerankers, change &lt;code&gt;Top-K&lt;/code&gt; limits, and modify prompt templates interactively.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Granular Execution Traces:&lt;/strong&gt; For every single trial, Muffakir records the full execution path. You get complete visibility into latency, cost, answer quality, retrieved context, and the exact queries generated by the LLM.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reproducible Trials:&lt;/strong&gt; Stop guessing. Muffakir saves configurations and checkpoints, ensuring you can replicate exact test conditions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Temporal Benchmarking Ready:&lt;/strong&gt; The architecture is designed to handle complex edge cases, such as evaluating how your RAG system reacts when source documents are superseded, corrected, or withdrawn over time.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  How It Works
&lt;/h2&gt;

&lt;p&gt;Getting started takes less than a minute. You can install the framework and spin up the &lt;strong&gt;ComposerUI&lt;/strong&gt; directly from your terminal:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Install Muffakir with standard dependencies&lt;/span&gt;
pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="s2"&gt;"Muffakir[standard]"&lt;/span&gt;

&lt;span class="c"&gt;# Launch the UI and start experimenting locally&lt;/span&gt;
muffakir serve &lt;span class="nt"&gt;--open&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;p&gt;If you are building LLM applications and want to stop flying blind when optimizing your RAG pipelines, I’d love for you to give it a spin!&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Check out the repo and give it a star ⭐️:&lt;/strong&gt; &lt;a href="https://github.com/Mohamed28112003/Muffakir" rel="noopener noreferrer"&gt;Mohamed28112003/Muffakir&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I'm actively looking for feedback, feature requests, and open-source contributors. What do you currently use to evaluate your RAG pipelines? Let me know in the comments!&lt;/p&gt;

</description>
      <category>ai</category>
      <category>rag</category>
      <category>opensource</category>
      <category>python</category>
    </item>
  </channel>
</rss>
