<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Robin Hayer</title>
    <description>The latest articles on DEV Community by Robin Hayer (@robinhayer).</description>
    <link>https://dev.to/robinhayer</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4070479%2F24b4cddd-eb98-4770-832f-dceeb57587b5.jpg</url>
      <title>DEV Community: Robin Hayer</title>
      <link>https://dev.to/robinhayer</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/robinhayer"/>
    <language>en</language>
    <item>
      <title>The tool said it wrote 48 packets. Only 44 were there</title>
      <dc:creator>Robin Hayer</dc:creator>
      <pubDate>Mon, 21 Sep 2026 00:58:20 +0000</pubDate>
      <link>https://dev.to/robinhayer/the-tool-said-it-wrote-48-packets-only-44-were-there-50k6</link>
      <guid>https://dev.to/robinhayer/the-tool-said-it-wrote-48-packets-only-44-were-there-50k6</guid>
      <description>&lt;p&gt;&lt;em&gt;Exit code zero said success, but the packet count said otherwise.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Recap
&lt;/h2&gt;

&lt;p&gt;I have been developing a tool wrapped around tshark. The &lt;a href="https://robinhayer.hashnode.dev/the-2-5-gb-wall" rel="noopener noreferrer"&gt;first blog post&lt;/a&gt; was about hitting a wall on a 2.5 GB file. Later, I talked about parallelizing the PCAP processing in my &lt;a href="https://robinhayer.hashnode.dev/concurrency-without-a-parallel-parser" rel="noopener noreferrer"&gt;second blog post&lt;/a&gt; where I ran into a file corruption bug.  &lt;/p&gt;

&lt;p&gt;Now here comes the real culprit: the PcapSplitter's &lt;code&gt;FiveTupleSplitter&lt;/code&gt; caused file truncation/corruption on TCP session reuse (i.e., a new&amp;nbsp;&lt;code&gt;SYN&lt;/code&gt; packet arrives for an already tracked 5-tuple hash).&lt;/p&gt;

&lt;h2&gt;
  
  
  How I found it
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://robinhayer.hashnode.dev/concurrency-without-a-parallel-parser" rel="noopener noreferrer"&gt;Post two&lt;/a&gt; ended with me replacing PcapSplitter. &lt;em&gt;This is what came before that replacement&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;The initial signal came from "&lt;strong&gt;Total Block Length&lt;/strong&gt;" errors thrown by tshark on some output files. That told me something was wrong, not what.&lt;/p&gt;

&lt;p&gt;The actual signal came when I built a small reproduction. PcapSplitter reported 12 files and 48 packets, but on disk, there were 11 files and 44 packets. Exit code zero and printed "Finished" on standard output.&lt;/p&gt;

&lt;h2&gt;
  
  
  Theory One: File Descriptor Exhaustion, wrong
&lt;/h2&gt;

&lt;p&gt;Someone on Reddit suggested file descriptor exhaustion. It was plausible with one output file per flow, where I had 95-125 flows, and a failed &lt;code&gt;open()&lt;/code&gt; with an unchecked return would produce silent drops.&lt;/p&gt;

&lt;p&gt;But instead of relying on theory, I tested it. Then I found that at low &lt;code&gt;ulimit -n&lt;/code&gt; it silently drops most packets and still exits zero. Sharp threshold, reproducible. It wasn't my corruption. Raising the limit didn't fix the original problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  Theory Two: Hardcoded LRU Limit, again wrong
&lt;/h2&gt;

&lt;p&gt;After testing file descriptor exhaustion, I found that the PcapSplitter library has a hardcoded &lt;code&gt;MAX_NUMBER_OF_CONCURRENT_OPEN_FILES = 250&lt;/code&gt; with an LRU that closes and reopens handles past it. It was a perfect candidate.&lt;/p&gt;

&lt;p&gt;I tested it too. It held up fine. My minimal reproduction was 13 connections, nowhere near the cap.&lt;/p&gt;

&lt;p&gt;Two good theories, one found a different real bug, the other held up fine. Neither was mine.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it was exactly
&lt;/h2&gt;

&lt;p&gt;The correct part of the splitter was assigning a new file number when a TCP session reuses a 5-tuple, but the filename function builds the name from IP and port only. So, in this case, both sessions get the same filenames. Then &lt;code&gt;main.cpp&lt;/code&gt; sees a file number it has never seen, and opens that file fresh, without append. This truncates the existing file or causes a race condition between two active file writer handlers.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. File Number Allocation on Session Reuse (&lt;code&gt;ConnectionSplitters.h&lt;/code&gt;)
&lt;/h3&gt;

&lt;p&gt;When a fresh&amp;nbsp;&lt;code&gt;SYN&lt;/code&gt; arrives for a 5-tuple hash that has been seen before,&amp;nbsp;&lt;code&gt;FiveTupleSplitter::getFileNumber&lt;/code&gt;&amp;nbsp;deliberately assigns a&amp;nbsp;&lt;strong&gt;new file number&lt;/strong&gt;&amp;nbsp;to isolate the logical session:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight cpp"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;isSyn&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="n"&gt;m_TcpFlowTable&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;find&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;hash&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="n"&gt;m_TcpFlowTable&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;end&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="n"&gt;m_TcpFlowTable&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;hash&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="nb"&gt;false&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;m_FlowTable&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;hash&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;getNextFileNumber&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;filesToClose&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;   &lt;span class="c1"&gt;// &amp;lt;-- Allocates a NEW file number&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  2. The Filename Disconnect (&lt;code&gt;ConnectionSplitters.h&lt;/code&gt;)
&lt;/h3&gt;

&lt;p&gt;Right below this logic,&amp;nbsp;&lt;code&gt;FiveTupleSplitter::getFileName&lt;/code&gt;&amp;nbsp;overrides the base class implementation but completely ignores the&amp;nbsp;&lt;code&gt;fileNumber&lt;/code&gt;&amp;nbsp;parameter passed into it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight cpp"&gt;&lt;code&gt;&lt;span class="n"&gt;std&lt;/span&gt;&lt;span class="o"&gt;::&lt;/span&gt;&lt;span class="n"&gt;string&lt;/span&gt; &lt;span class="nf"&gt;getFileName&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;pcpp&lt;/span&gt;&lt;span class="o"&gt;::&lt;/span&gt;&lt;span class="n"&gt;Packet&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt; &lt;span class="n"&gt;packet&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;const&lt;/span&gt; &lt;span class="n"&gt;std&lt;/span&gt;&lt;span class="o"&gt;::&lt;/span&gt;&lt;span class="n"&gt;string&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt; &lt;span class="n"&gt;outputPcapBasePath&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;fileNumber&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="c1"&gt;// ...&lt;/span&gt;
    &lt;span class="n"&gt;sstream&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&amp;lt;&lt;/span&gt; &lt;span class="s"&gt;"connection-"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;packet&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;isPacketOfType&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;pcpp&lt;/span&gt;&lt;span class="o"&gt;::&lt;/span&gt;&lt;span class="n"&gt;TCP&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="c1"&gt;// ...&lt;/span&gt;
        &lt;span class="c1"&gt;// Filename is derived strictly from the packet's IP/Port values:&lt;/span&gt;
        &lt;span class="n"&gt;updateStringStream&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sstream&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;getSrcIPString&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;packet&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;srcPort&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;getDstIPString&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;packet&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;dstPort&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;outputPcapBasePath&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;sstream&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;str&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;   &lt;span class="c1"&gt;// &amp;lt;-- fileNumber parameter is UNUSED on the TCP/UDP path&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="c1"&gt;// ...&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Because&amp;nbsp;&lt;code&gt;fileNumber&lt;/code&gt;&amp;nbsp;is dropped, the new session produces the&amp;nbsp;&lt;strong&gt;identical filename string&lt;/strong&gt;&amp;nbsp;as the older session.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. The Truncation / Race Condition (&lt;code&gt;main.cpp&lt;/code&gt;)
&lt;/h3&gt;

&lt;p&gt;In&amp;nbsp;&lt;code&gt;main.cpp&lt;/code&gt;, the active writer cache map (&lt;code&gt;outputFiles&lt;/code&gt;) is keyed by the integer&amp;nbsp;&lt;code&gt;fileNum&lt;/code&gt;, not the filepath string. When the newly allocated&amp;nbsp;&lt;code&gt;fileNum&lt;/code&gt;&amp;nbsp;is evaluated, it goes down the fresh initialization branch:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight cpp"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Since fileNum is a newly generated integer, it won't be found in the map&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;outputFiles&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;find&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;fileNum&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;outputFiles&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;end&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt; 
&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;std&lt;/span&gt;&lt;span class="o"&gt;::&lt;/span&gt;&lt;span class="n"&gt;string&lt;/span&gt; &lt;span class="n"&gt;fileName&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;splitter&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="n"&gt;getFileName&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;parsedPacket&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;outputPcapFileName&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;fileNum&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;outputFileExtension&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;isReaderPcapng&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;outputFiles&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;fileNum&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;reset&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="n"&gt;pcpp&lt;/span&gt;&lt;span class="o"&gt;::&lt;/span&gt;&lt;span class="n"&gt;PcapNgFileWriterDevice&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;fileName&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;outputFiles&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;fileNum&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;reset&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="n"&gt;pcpp&lt;/span&gt;&lt;span class="o"&gt;::&lt;/span&gt;&lt;span class="n"&gt;PcapFileWriterDevice&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;fileName&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;rawPacket&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;getLinkLayerType&lt;/span&gt;&lt;span class="p"&gt;()));&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="c1"&gt;// CRITICAL: Plain open — no append flag — on a path that already holds valid session data!&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="n"&gt;outputFiles&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;fileNum&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="n"&gt;open&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt; 
        &lt;span class="k"&gt;break&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Impact &amp;amp; Consequences
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;File Truncation:&lt;/strong&gt;&amp;nbsp;If the previous file representing the same session was closed by the LRU mechanism, the new file writer triggers a standard file open on the existing path,&amp;nbsp;that truncates the existing PCAP data captured from the previous session.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Write Race / Corruption:&lt;/strong&gt;&amp;nbsp;If duplicate files representing the same session are still open and active, it creates a race condition between two active file writer handlers. This breaks the linear layout of the PCAP/PCAPNG format and leaves the file corrupted.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Reporting the bug and fixing upstream
&lt;/h2&gt;

&lt;p&gt;I filed the issue with a 13-connection reproduction in &lt;a href="https://github.com/seladb/PcapPlusPlus/issues/2248" rel="noopener noreferrer"&gt;#2248&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The maintainer found his own commit &lt;a href="https://github.com/seladb/PcapPlusPlus/commit/63791006db94bc07aeb88f55a649d80dc121cbc2" rel="noopener noreferrer"&gt;6379100&lt;/a&gt; and said he couldn't remember why he'd made the change. He pushed back on my first fix, where I suggested appending the number to every filename. But it changes output for everyone, which is why he pushed back. Fair point.&lt;/p&gt;

&lt;p&gt;Then I proposed a middle ground where we only suffix on actual collision, first session keeps its name. He accepted it, asked for tests. I wrote the code and tests.&lt;/p&gt;

&lt;p&gt;Merged &lt;a href="https://github.com/seladb/PcapPlusPlus/pull/2249" rel="noopener noreferrer"&gt;#2249&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F94iottchgai94vg6y4a8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F94iottchgai94vg6y4a8.png" alt="The maintainer approving and merging the fix, with a thank-you comment." width="800" height="563"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Takeaways
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Exit code zero is not always success. Verify what a tool claims against what it produced. &lt;/li&gt;
&lt;li&gt;A possible theory that reproduces a real bug can still be the wrong theory.&lt;/li&gt;
&lt;li&gt;Bugs usually come from two reasonable components disagreeing about a definition. Not necessarily a broken piece.&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>debugging</category>
      <category>opensource</category>
      <category>cpp</category>
      <category>networking</category>
    </item>
    <item>
      <title>Concurrency Without a Parallel Parser: Splitting PCAPs by Session</title>
      <dc:creator>Robin Hayer</dc:creator>
      <pubDate>Wed, 26 Aug 2026 10:00:00 +0000</pubDate>
      <link>https://dev.to/robinhayer/concurrency-without-a-parallel-parser-splitting-pcaps-by-session-1op4</link>
      <guid>https://dev.to/robinhayer/concurrency-without-a-parallel-parser-splitting-pcaps-by-session-1op4</guid>
      <description>&lt;p&gt;&lt;em&gt;tshark needs linear state, so I moved the concurrency upstream — and hit a corrupted-output problem on the way.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the last post left off
&lt;/h2&gt;

&lt;p&gt;In my first post about &lt;a href="https://robinhayer.dev/the-2-5-gb-wall" rel="noopener noreferrer"&gt;the 2.5 GB wall&lt;/a&gt;, streaming fixed memory. It didn't fix throughput.&lt;/p&gt;

&lt;p&gt;The pipeline was still one file, one tshark process, one core. Piping stdout to stdin doesn't parallelize anything — it just stops the pipeline from holding the whole file in memory at once. Processing time dropped, but the shape of the work didn't change: still sequential, still bottlenecked on a single core no matter how large the machine underneath it was.&lt;/p&gt;

&lt;p&gt;A few people asked the obvious question in the comments: why not just add goroutines?&lt;/p&gt;

&lt;p&gt;I tried. It didn't work, and the reason it didn't work is more interesting than the fix.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the parser can't parallelize
&lt;/h2&gt;

&lt;p&gt;Goroutines help when work can be split into independent pieces. tshark's dissection isn't independent — it's linear state.&lt;/p&gt;

&lt;p&gt;It reads a capture packet by packet, and what it reads in one packet changes how it decodes the next. TCP stream reassembly needs the packets in order to rebuild the byte stream. Connection state tracks which flow a packet belongs to. Dissectors that depend on earlier context — retransmission detection, duplicate ACK detection, anything under &lt;code&gt;tcp.analysis.*&lt;/code&gt; — read and update shared conversation tables as they go.&lt;/p&gt;

&lt;p&gt;None of that is optional. tshark wasn't written single-threaded because nobody got around to parallelizing it — the dissection model requires strict ordering. Skip packets out of order and the state tshark is tracking for that stream goes wrong.&lt;/p&gt;

&lt;p&gt;So wrapping the consuming side in goroutines doesn't help, because the problem was never on the consuming side. My Go program reading tshark's output line by line was already fast. The producing side — tshark itself, walking the file front to back — is sequential no matter what I do to the code around it.&lt;/p&gt;

&lt;p&gt;If I wanted concurrency, it had to happen before tshark ever saw the file.&lt;/p&gt;

&lt;h2&gt;
  
  
  Move the concurrency upstream
&lt;/h2&gt;

&lt;p&gt;If one tshark process can't go faster, run several over pieces of the file at once.&lt;/p&gt;

&lt;p&gt;The obvious version of that is wrong. You can't cut a pcap into equal-sized chunks — a TCP stream that straddles a cut point loses the packets on the other side of the boundary, and the dissector on that chunk has no idea the connection existed before the cut. Retransmission flags, sequence tracking, anything stateful breaks the moment a conversation is split across two files.&lt;/p&gt;

&lt;p&gt;The fix is to split along conversation boundaries instead of byte boundaries. Group packets by session — same 5-tuple, same connection — so every chunk holds complete conversations and nothing crosses a chunk boundary mid-stream. No state is lost, because no state ever needed to survive past a single chunk.&lt;/p&gt;

&lt;p&gt;Once the split is done that way, the chunks are independent by construction. Each one goes to its own tshark process, running in parallel, merged back together at the end. The dissection itself stays linear — it just becomes N linear passes running at once instead of one linear pass running alone.&lt;/p&gt;

&lt;p&gt;That was the plan. Getting there took one extra tool falling apart on me.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffrarwmfl0jydtn3nxf3y.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffrarwmfl0jydtn3nxf3y.png" alt="Comparison of two splitting strategies. Splitting by size cuts conversation A across two chunks, losing state at the boundary. Splitting by session keeps each conversation complete within a single chunk." width="799" height="348"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The bug
&lt;/h2&gt;

&lt;p&gt;I reached for PcapSplitter, part of PcapPlusPlus, running in connection mode — split by 5-tuple, exactly what I needed.&lt;/p&gt;

&lt;p&gt;In connection mode it holds one output file open per distinct flow for the length of the pass. On the files that mattered, that was 95 to 125 distinct flows open at once. Under that load, the output came back corrupted.&lt;/p&gt;

&lt;p&gt;Two distinct failure signatures:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;total block length N of an EPB is too small for M bytes of packet data&lt;/li&gt;
&lt;li&gt;total block length 0 of an unknown block type is less than the minimum block size 12&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I reproduced both on the master build and on the v25.05 stable release, so this wasn't something already fixed between versions with me sitting on a stale checkout. I reproduced both on pcapng and on legacy pcap after converting with &lt;code&gt;editcap -F pcap&lt;/code&gt;, so it wasn't specific to the pcapng container format either. And I ran the source file through &lt;code&gt;pcapfix&lt;/code&gt; before assuming the input was the problem — it came back clean. The corruption was being introduced by the split, not inherited from the source.&lt;/p&gt;

&lt;p&gt;Same root cause both times: too many output file handles held open at once by one process. I didn't dig further into PcapSplitter's internals to find the exact line — I didn't need to, because the fix wasn't going to be a patch. I replaced it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reimplementing it in-process
&lt;/h2&gt;

&lt;p&gt;I rewrote the session split using gopacket, inside my own program, instead of shelling out to a separate tool. No external process, no file handles held open by something I don't control — the splitting and the writing happen where I can see and bound exactly what's open at once.&lt;/p&gt;

&lt;p&gt;Ran it 20 times across both files that had triggered the original corruption, in both the Go and Python versions of the pipeline. 20 out of 20 succeeded, including full concurrency on the file that had reproducibly failed before.&lt;/p&gt;

&lt;h2&gt;
  
  
  The honest ending
&lt;/h2&gt;

&lt;p&gt;Here's the part worth sitting with.&lt;/p&gt;

&lt;p&gt;Session splitting only kicks in above 100,000 packets — below that, one tshark process handles the whole file fine. I went back through the real corpus this pipeline processes day to day: 57 files. Exactly 3 of them cross that threshold.&lt;/p&gt;

&lt;p&gt;I built session splitting, and everything downstream of it, for a problem that shows up in about 1 out of every 19 files. It was the right thing to build — the 3 files that need it would otherwise take hours or fail outright — but it's worth saying plainly: most of the corpus never touches this code path at all.&lt;/p&gt;

</description>
      <category>go</category>
      <category>networking</category>
      <category>performance</category>
      <category>wireshark</category>
    </item>
    <item>
      <title>The 2.5 GB Wall: Why My tshark Wrapper Died, and How Streaming Fixed It</title>
      <dc:creator>Robin Hayer</dc:creator>
      <pubDate>Mon, 10 Aug 2026 04:22:47 +0000</pubDate>
      <link>https://dev.to/robinhayer/the-25-gb-wall-why-my-tshark-wrapper-died-and-how-streaming-fixed-it-3nfg</link>
      <guid>https://dev.to/robinhayer/the-25-gb-wall-why-my-tshark-wrapper-died-and-how-streaming-fixed-it-3nfg</guid>
      <description>&lt;h2&gt;
  
  
  The tool that worked
&lt;/h2&gt;

&lt;p&gt;I built a CLI tool that wrapped tshark for PCAP analysis. Grab the file, send shell commands, let tshark analyze it, capture the output, shape it, return it. Quite simple, right.&lt;/p&gt;

&lt;p&gt;It was a good simple design, and it was the right design for what I was doing. I was analyzing small PCAP files and it worked smoothly. No hidden bugs, no errors, nothing to fix. When something works, you stop thinking about it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Then came 2.5 GB
&lt;/h2&gt;

&lt;p&gt;It was fine until I hit a 2.5 GB file. Six to seven hours of processing, and repeated crashes with out-of-memory errors.&lt;/p&gt;

&lt;p&gt;I had never seen anything like this before. I ran it again. Same thing. Hours in diagnosis, still nothing. My first instinct was that I was messing something up in the shell commands, or that something was wrong with tshark itself.&lt;/p&gt;

&lt;p&gt;Honestly, I was happy about it. I had never been challenged like this before. So I took the challenge.&lt;/p&gt;

&lt;p&gt;The one thing I was sure of was the scale: this file held around 1.9 million packets. The problem was either on tshark's side or mine. It turned out to be both.&lt;/p&gt;

&lt;h2&gt;
  
  
  What was actually wrong
&lt;/h2&gt;

&lt;p&gt;Three things, and they looked like one.&lt;/p&gt;

&lt;p&gt;First, I was running three separate queries against the same file — one for analytics, one for rows, one for full dissection. Three full passes over 2.5 GB. I had written them as three because they answered three different questions, and at small file sizes that cost nothing. At 2.5 GB it cost everything.&lt;/p&gt;

&lt;p&gt;Second, tshark is single-threaded. For a capture this size, there is no parallelism to fall back on.&lt;/p&gt;

&lt;p&gt;Third — and this was the part I was most responsible for — I was asking for full JSON dissection. tshark built the entire output in memory, and then my program parsed all of it in memory. Two full copies of a dataset that was already too large.&lt;/p&gt;

&lt;p&gt;I couldn't go further with that architecture. If I don't have memory, I can't process it. No amount of tuning fixes a design that requires holding everything at once.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffosmx6xtnp1i2533ufc1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffosmx6xtnp1i2533ufc1.png" alt="Pipeline showing a 2.5 GB PCAP file passed to tshark for full JSON dissection, with the full output held in memory and parsed entirely in memory before producing output. The two memory-holding stages are highlighted in red." width="800" height="69"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Cutting three passes down to one
&lt;/h2&gt;

&lt;p&gt;The first fix was the obvious one once I saw it: stop reading the file three times.&lt;/p&gt;

&lt;p&gt;I consolidated the three queries into a single optimized query that returned everything I needed in one pass. Same output, one traversal instead of three.&lt;/p&gt;

&lt;p&gt;That alone took processing from six to seven hours down to one to two hours. No architecture change, no extra hardware. Just not doing the same expensive work three times.&lt;/p&gt;

&lt;p&gt;It was a big win and it wasn't enough. The memory pressure was still there, and one to two hours for a single file is still a bad number.&lt;/p&gt;

&lt;h2&gt;
  
  
  Streaming the output
&lt;/h2&gt;

&lt;p&gt;The second fix came from a concept I knew in theory but had never needed: streaming.&lt;/p&gt;

&lt;p&gt;Instead of waiting for tshark to finish and hand me a complete result, I connected tshark's stdout directly to my Go program's stdin. Now tshark emits packets and my program consumes them as they arrive. Processing overlaps generation. Memory stays flat, because at any moment I'm only holding one packet, not 1.9 million.&lt;/p&gt;

&lt;p&gt;Then I narrowed the query itself. I stopped asking for full JSON dissection and switched to &lt;code&gt;-T fields&lt;/code&gt; and &lt;code&gt;-T ek&lt;/code&gt;, requesting only the fields I actually needed. Full JSON dissection means tshark decodes every layer of every packet and serializes all of it. Most of that work was thrown away by my program a moment later. Narrowing the query removed the work instead of optimizing it.&lt;/p&gt;

&lt;p&gt;Together those two changes rebuilt the pipeline: tshark emits selected fields packet by packet, my Go program reads line by line, processes, and writes output as it goes.&lt;/p&gt;

&lt;p&gt;The OOM crashes were gone instantly. Processing dropped to about 70 minutes on the same file — down from six to seven hours originally, on the same machine, with no extra cores.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3o6rc5d1dcnohwjyb2dh.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3o6rc5d1dcnohwjyb2dh.png" alt="Pipeline showing a 2.5 GB PCAP file passed to tshark with -T fields returning only selected fields, its stdout piped to stdin of a Go reader processing line by line at constant memory, producing an output stream. The tshark and Go reader stages are highlighted in green." width="798" height="76"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What it didn't fix
&lt;/h2&gt;

&lt;p&gt;Streaming solved memory. It did not solve throughput.&lt;/p&gt;

&lt;p&gt;The operation was still single-threaded. Nothing about piping stdout to stdin makes tshark use more cores. And if multiple files arrive at once, the problem comes straight back — one file at a time, one core, everything else waiting.&lt;/p&gt;

&lt;p&gt;That's the wall I hit next, and it's a different problem with a different fix.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;If you've hit something similar, I'd like to hear how you handled it. I'm writing up the concurrency side next.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>go</category>
      <category>performance</category>
      <category>networking</category>
      <category>wireshark</category>
    </item>
  </channel>
</rss>
