<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Karol Kurzydym</title>
    <description>The latest articles on DEV Community by Karol Kurzydym (@karol_kurzydym_4f36580c50).</description>
    <link>https://dev.to/karol_kurzydym_4f36580c50</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4162299%2F4bafaeb2-85ad-4475-ae0a-5f779684e71e.png</url>
      <title>DEV Community: Karol Kurzydym</title>
      <link>https://dev.to/karol_kurzydym_4f36580c50</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/karol_kurzydym_4f36580c50"/>
    <language>en</language>
    <item>
      <title>Three tiny CSV fixtures that catch different importer mistakes</title>
      <dc:creator>Karol Kurzydym</dc:creator>
      <pubDate>Sun, 04 Oct 2026 18:41:26 +0000</pubDate>
      <link>https://dev.to/karol_kurzydym_4f36580c50/three-tiny-csv-fixtures-that-catch-different-importer-mistakes-cha</link>
      <guid>https://dev.to/karol_kurzydym_4f36580c50/three-tiny-csv-fixtures-that-catch-different-importer-mistakes-cha</guid>
      <description>&lt;p&gt;A well-formed CSV can still carry the wrong meaning for your importer. Here are&lt;br&gt;
three fictional inputs, with explicit expectations, that you can copy into a test&lt;br&gt;
suite without uploading any customer data.&lt;/p&gt;
&lt;h3&gt;
  
  
  An identifier is a string
&lt;/h3&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;id,postal_code
0007,00123
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Expected cells: &lt;code&gt;["0007", "00123"]&lt;/code&gt;. If your application converts both values to&lt;br&gt;
integers, it has changed the identifiers. Whether to convert a column belongs in&lt;br&gt;
the application's schema, not in a guess based on the characters in one record.&lt;/p&gt;
&lt;h3&gt;
  
  
  A comma inside quotes is one cell
&lt;/h3&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;id,name
001,"Kowalski, Jan"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Expected: one data record, two cells. Splitting a line on commas produces three&lt;br&gt;
cells instead. Use a CSV parser. The same testing principle applies to escaped&lt;br&gt;
quotes and multiline fields: physical line count is not necessarily record count.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;csv&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;io&lt;/span&gt;
&lt;span class="n"&gt;text&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;id,name&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s"&gt;001,&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Kowalski, Jan&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;
&lt;span class="n"&gt;rows&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;list&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;csv&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;reader&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;io&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;StringIO&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;newline&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;""&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;strict&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;rows&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="p"&gt;[[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;001&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Kowalski, Jan&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Repeated IDs need a declared rule
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;id,name
001,Ada
001,Jan
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both records parse successfully. An importer that keeps the last row silently has&lt;br&gt;
made a business decision. A useful regression test states that decision explicitly:&lt;br&gt;
reject the conflict, flag it for review, or resolve it using a documented rule.&lt;br&gt;
For a neutral diagnostic preview, preserving both rows and flagging the conflict&lt;br&gt;
is easier to review than deleting either record.&lt;/p&gt;

&lt;p&gt;Record the input's encoding and delimiter beside each fixture. Test exact parsed&lt;br&gt;
cells, not only a success flag. Separate parser failures from application warnings.&lt;br&gt;
A duplicate record need not be a malformed CSV, and a valid CSV need not be an&lt;br&gt;
acceptable input for your particular import.&lt;/p&gt;

&lt;p&gt;Source for parser behavior: &lt;a href="https://docs.python.org/3/library/csv.html" rel="noopener noreferrer"&gt;https://docs.python.org/3/library/csv.html&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Disclosure: these examples and explanatory text were prepared with AI assistance.&lt;br&gt;
The data are fictional. The parser example was executed on Python 3.12 in Linux.&lt;br&gt;
Which additional input has caught an importer bug in your test suite? I am collecting concrete missing cases before expanding these examples.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Optional paid pack:&lt;/strong&gt; I also prepared a separate pack of 24 fictional CSV fixtures, an exact-results manifest and a Python verifier. The proposed price is &lt;strong&gt;49 PLN&lt;/strong&gt; for a license to use it in your internal tests. If you want the contents and purchase terms, ask in the comments; this is my own product, not an independent recommendation. There is no checkout link in this post. Please do not post customer data or payment information.&lt;/p&gt;

</description>
      <category>testing</category>
      <category>python</category>
    </item>
  </channel>
</rss>
