<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Ivan</title>
    <description>The latest articles on DEV Community by Ivan (@ijne).</description>
    <link>https://dev.to/ijne</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4149432%2F75433979-d143-411d-80b2-f3cedc854e2c.png</url>
      <title>DEV Community: Ivan</title>
      <link>https://dev.to/ijne</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ijne"/>
    <language>en</language>
    <item>
      <title>Built a local AI pipeline that captures system audio, transcribes it, structures the result with a local LLM, and sends reviewed notes to Obsidian. I also wrote about what broke when real users started testing it.</title>
      <dc:creator>Ivan</dc:creator>
      <pubDate>Tue, 29 Sep 2026 18:20:58 +0000</pubDate>
      <link>https://dev.to/ijne/built-a-local-ai-pipeline-that-captures-system-audio-transcribes-it-structures-the-result-with-a-gmj</link>
      <guid>https://dev.to/ijne/built-a-local-ai-pipeline-that-captures-system-audio-transcribes-it-structures-the-result-with-a-gmj</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/ijne/building-memplua-a-local-ai-pipeline-for-turning-system-audio-into-obsidian-notes-pb6" class="crayons-story__hidden-navigation-link"&gt;Building Memplua: A Local AI Pipeline for Turning System Audio Into Obsidian Notes&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/ijne" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4149432%2F75433979-d143-411d-80b2-f3cedc854e2c.png" alt="ijne profile" class="crayons-avatar__image"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/ijne" class="crayons-story__secondary fw-medium m:hidden"&gt;
              Ivan
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                Ivan
                
                
              
              &lt;div id="story-author-preview-content-4771067" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/ijne" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4149432%2F75433979-d143-411d-80b2-f3cedc854e2c.png" class="crayons-avatar__image" alt=""&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;Ivan&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/ijne/building-memplua-a-local-ai-pipeline-for-turning-system-audio-into-obsidian-notes-pb6" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Sep 29&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/ijne/building-memplua-a-local-ai-pipeline-for-turning-system-audio-into-obsidian-notes-pb6" id="article-link-4771067"&gt;
          Building Memplua: A Local AI Pipeline for Turning System Audio Into Obsidian Notes
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/ai"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;ai&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/opensource"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;opensource&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/education"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;education&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/softwareengineering"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;softwareengineering&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
          &lt;a href="https://dev.to/ijne/building-memplua-a-local-ai-pipeline-for-turning-system-audio-into-obsidian-notes-pb6" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left"&gt;
            &lt;div class="multiple_reactions_aggregate"&gt;
              &lt;span class="multiple_reactions_icons_container"&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/sparkle-heart-5f9bee3767e18deb1bb725290cb151c25234768a0e9a2bd39370c382d02920cf.svg" width="18" height="18"&gt;
                  &lt;/span&gt;
              &lt;/span&gt;
              &lt;span class="aggregate_reactions_counter"&gt;1&lt;span class="hidden s:inline"&gt;&amp;nbsp;reaction&lt;/span&gt;&lt;/span&gt;
            &lt;/div&gt;
          &lt;/a&gt;
            &lt;a href="https://dev.to/ijne/building-memplua-a-local-ai-pipeline-for-turning-system-audio-into-obsidian-notes-pb6#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              &lt;span class="hidden s:inline"&gt;Add&amp;nbsp;Comment&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            5 min read
          &lt;/small&gt;
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


</description>
      <category>ai</category>
      <category>automation</category>
      <category>llm</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Building Memplua: A Local AI Pipeline for Turning System Audio Into Obsidian Notes</title>
      <dc:creator>Ivan</dc:creator>
      <pubDate>Tue, 29 Sep 2026 11:25:05 +0000</pubDate>
      <link>https://dev.to/ijne/building-memplua-a-local-ai-pipeline-for-turning-system-audio-into-obsidian-notes-pb6</link>
      <guid>https://dev.to/ijne/building-memplua-a-local-ai-pipeline-for-turning-system-audio-into-obsidian-notes-pb6</guid>
      <description>&lt;p&gt;I started building Memplua because I wanted my knowledge base to fill itself. Not literally, of course. The real problem was the gap between consuming information and turning it into something structured. I could watch a technical video, listen to a podcast, or hear something useful during work, but if I wanted that information in Obsidian, I still had to stop, extract the important parts, rewrite them, and organize everything manually.&lt;/p&gt;

&lt;p&gt;That was the part I wanted to automate.&lt;br&gt;
The idea behind Memplua is simple: information already passes through your computer, so instead of asking the user to upload every source manually, the application should capture that stream, process it locally, and prepare structured notes for review.&lt;br&gt;
The project is open source and currently focused on Windows.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pipeline
&lt;/h2&gt;

&lt;p&gt;The current audio pipeline looks roughly like this:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;system audio&lt;/li&gt;
&lt;li&gt;voice activity detection&lt;/li&gt;
&lt;li&gt;transcription&lt;/li&gt;
&lt;li&gt;text chunking&lt;/li&gt;
&lt;li&gt;local LLM&lt;/li&gt;
&lt;li&gt;structured notes&lt;/li&gt;
&lt;li&gt;manual review&lt;/li&gt;
&lt;li&gt;Obsidian&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The main design goal was to keep the system generic. I did not want to build separate integrations for YouTube, browsers, Discord, Telegram, and every other possible source. If the information is already being played through the operating system, system audio gives me one universal input layer.&lt;/p&gt;

&lt;p&gt;For the same reason, I currently see Memplua more as a capture-and-interpretation pipeline than as a collection of service-specific connectors.&lt;/p&gt;

&lt;h2&gt;
  
  
  Capturing audio on Windows
&lt;/h2&gt;

&lt;p&gt;I started with audio because it seemed like the most difficult and useful input type.&lt;/p&gt;

&lt;p&gt;System audio capture is highly OS-specific, so I chose Windows first. It is also the platform I use myself, which made debugging much easier.&lt;br&gt;
The capture layer produces raw audio samples. These samples then go through preprocessing and are eventually passed to transcription.&lt;/p&gt;

&lt;p&gt;I use Whisper for speech-to-text because sending raw audio directly into a general-purpose LLM would be both inefficient and unnecessary.&lt;br&gt;
Once transcription is working, the problem changes from audio processing to context management.&lt;/p&gt;

&lt;h2&gt;
  
  
  The chunking problem
&lt;/h2&gt;

&lt;p&gt;Speech transcription arrives as chunks, but language does not care about chunk boundaries.&lt;br&gt;
A phrase like: Docker is a tool may be separated from: for containerization. If the LLM processes both independently, the second fragment becomes almost meaningless.&lt;/p&gt;

&lt;p&gt;My first implementation used a floating window: the next group of text reused part of the previous group. This helped preserve context, but it created another problem. The model started seeing the same information several times and generated duplicated notes.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F52wowz75dl4pqt4d10fr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F52wowz75dl4pqt4d10fr.png" alt=" " width="800" height="308"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This became especially annoying after I added manual review, because users had to approve multiple versions of the same term or idea.&lt;/p&gt;

&lt;p&gt;I eventually changed the strategy and separated the input into two parts:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;focus — text from which the model is allowed to extract new information;&lt;/li&gt;
&lt;li&gt;context — previous text that can only help interpret the focus section.
This sounds like a small change, but it made the pipeline much cleaner. The model can still understand incomplete phrases using previous context without continuously generating entities from the same old text.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frd3tnqzl5he75d4nmxrd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frd3tnqzl5he75d4nmxrd.png" alt=" " width="487" height="220"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Turning text into structured knowledge
&lt;/h2&gt;

&lt;p&gt;The LLM does not return free-form prose. I ask it to produce a structured JSON response that can be parsed by the application. Internally, I currently work with a few simple concepts.&lt;br&gt;
A term is a self-contained entity with a name and definition. A thought connects multiple terms. A group of extracted terms and thoughts becomes a draft note set that can later be reviewed by the user.&lt;br&gt;
For local inference I use llama.cpp, and one of the models I have been testing is Qwen3-4B-Q4.&lt;/p&gt;

&lt;p&gt;The important part for me is that inference happens locally. Memplua may process conversations, technical content, personal audio, or other information that I do not want to send to external APIs.&lt;br&gt;
Local inference also keeps the project usable without subscriptions, although it introduces another problem: setup becomes more complicated.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why I stopped fully automatic writing
&lt;/h2&gt;

&lt;p&gt;My first mental model was very simple:&lt;br&gt;
audio → LLM → Obsidian&lt;/p&gt;

&lt;p&gt;In practice, I quickly realized that letting an LLM write directly into a knowledge base is a terrible default.The model can misunderstand something, duplicate information, create bad relationships, or simply produce something that is technically correct but useless.&lt;/p&gt;

&lt;p&gt;So I added a mandatory review step. The application prepares a draft, but the user can edit or reject it before anything is saved. This makes the system less magical, but much more trustworthy.&lt;/p&gt;

&lt;p&gt;It also changes the purpose of the tool. Memplua is not supposed to replace thinking. It is supposed to remove the repetitive work required to turn consumed information into a usable draft.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4crvyos1clmhthh35ftr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4crvyos1clmhthh35ftr.png" alt=" " width="800" height="516"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  VAD became more difficult than expected
&lt;/h2&gt;

&lt;p&gt;I added Silero VAD to remove silence and noise before transcription.&lt;br&gt;
In theory, this should be straightforward: detect whether a chunk contains speech and avoid wasting resources on empty audio.&lt;/p&gt;

&lt;p&gt;In practice, I started getting unexpectedly low probabilities even on obvious human speech.&lt;/p&gt;

&lt;p&gt;One external tester later reproduced the same behaviour. Their microphone signal had perfectly reasonable amplitude values:&lt;br&gt;
&lt;code&gt;peak=0.55&lt;br&gt;
rms=0.24&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;but Silero returned:&lt;br&gt;
&lt;code&gt;max_vad_probability=0.0013&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;The audio was then dropped as if no voice was present. When VAD was disabled, transcription and note generation started working normally.&lt;br&gt;
That narrowed the issue down significantly. The problem is likely somewhere in the VAD input pipeline: sample rate, normalization, window size, tensor shape, state handling, or a combination of those.&lt;/p&gt;

&lt;p&gt;For now, this is one of the parts I still consider unfinished.&lt;/p&gt;

&lt;h2&gt;
  
  
  Then real users started testing it
&lt;/h2&gt;

&lt;p&gt;After I published Memplua, external users immediately found problems that I had missed.&lt;/p&gt;

&lt;p&gt;One build was missing a runtime DLL because my own development machine already had the dependency installed. Another bug caused the process to launch while the UI itself remained invisible.&lt;/p&gt;

&lt;p&gt;Both problems are now fixed, and I rebuilt the installer so dependencies are handled more reliably. I also added both online and offline installation paths.&lt;/p&gt;

&lt;p&gt;The more interesting part was not the bugs themselves, but the quality of the feedback. People started sending logs, testing different configurations, opening GitHub issues, and even helping each other inside issue discussions.&lt;/p&gt;

&lt;p&gt;For a small project, that was probably the first moment when Memplua started to feel like an actual open-source project instead of just my personal repository.&lt;/p&gt;

&lt;h2&gt;
  
  
  Current state
&lt;/h2&gt;

&lt;p&gt;Memplua is still alpha software, but the main pipeline already works.&lt;br&gt;
At the moment, the project has:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Windows system audio capture;&lt;/li&gt;
&lt;li&gt;transcription;&lt;/li&gt;
&lt;li&gt;local LLM inference;&lt;/li&gt;
&lt;li&gt;structured note generation;&lt;/li&gt;
&lt;li&gt;manual review and editing;&lt;/li&gt;
&lt;li&gt;Obsidian export;&lt;/li&gt;
&lt;li&gt;online and offline installers;&lt;/li&gt;
&lt;li&gt;fixed UI startup behaviour;&lt;/li&gt;
&lt;li&gt;early external testers and active issues.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The next work is mostly about reliability and usability rather than adding more AI features. I want first-run setup to be easier, VAD behaviour to be more predictable, model configuration to be clearer, and local inference presets to require less manual tuning. That is less exciting than adding another model, but probably much more important.&lt;/p&gt;

&lt;h2&gt;
  
  
  Final thoughts
&lt;/h2&gt;

&lt;p&gt;The most useful thing about releasing Memplua was not getting repository views or stars. It was watching people use the application differently from how I use it. They found bugs I had stopped noticing, questioned UI decisions that seemed obvious to me, and forced me to think more carefully about the boundaries of the project.&lt;/p&gt;

&lt;p&gt;That feedback has already changed the application more than another week of development in isolation probably would have.&lt;br&gt;
If you are interested in local AI, Windows audio processing, Obsidian, or local inference, feedback is welcome.&lt;br&gt;
GitHub: &lt;a href="https://github.com/Ijne/memplua" rel="noopener noreferrer"&gt;https://github.com/Ijne/memplua&lt;/a&gt;&lt;br&gt;
Website: &lt;a href="https://memplua.space" rel="noopener noreferrer"&gt;https://memplua.space&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>education</category>
      <category>softwareengineering</category>
    </item>
  </channel>
</rss>
