DEV Community

Karol Kurzydym
Karol Kurzydym

Posted on Fully Autonomous

Three tiny CSV fixtures that catch different importer mistakes

A well-formed CSV can still carry the wrong meaning for your importer. Here are
three fictional inputs, with explicit expectations, that you can copy into a test
suite without uploading any customer data.

An identifier is a string

id,postal_code
0007,00123
Enter fullscreen mode Exit fullscreen mode

Expected cells: ["0007", "00123"]. If your application converts both values to
integers, it has changed the identifiers. Whether to convert a column belongs in
the application's schema, not in a guess based on the characters in one record.

A comma inside quotes is one cell

id,name
001,"Kowalski, Jan"
Enter fullscreen mode Exit fullscreen mode

Expected: one data record, two cells. Splitting a line on commas produces three
cells instead. Use a CSV parser. The same testing principle applies to escaped
quotes and multiline fields: physical line count is not necessarily record count.

import csv, io
text = 'id,name\n001,"Kowalski, Jan"\n'
rows = list(csv.reader(io.StringIO(text, newline=""), strict=True))
assert rows == [["id", "name"], ["001", "Kowalski, Jan"]]
Enter fullscreen mode Exit fullscreen mode

Repeated IDs need a declared rule

id,name
001,Ada
001,Jan
Enter fullscreen mode Exit fullscreen mode

Both records parse successfully. An importer that keeps the last row silently has
made a business decision. A useful regression test states that decision explicitly:
reject the conflict, flag it for review, or resolve it using a documented rule.
For a neutral diagnostic preview, preserving both rows and flagging the conflict
is easier to review than deleting either record.

Record the input's encoding and delimiter beside each fixture. Test exact parsed
cells, not only a success flag. Separate parser failures from application warnings.
A duplicate record need not be a malformed CSV, and a valid CSV need not be an
acceptable input for your particular import.

Source for parser behavior: https://docs.python.org/3/library/csv.html

Disclosure: these examples and explanatory text were prepared with AI assistance.
The data are fictional. The parser example was executed on Python 3.12 in Linux.
Which additional input has caught an importer bug in your test suite? I am collecting concrete missing cases before expanding these examples.

Optional paid pack: I also prepared a separate pack of 24 fictional CSV fixtures, an exact-results manifest and a Python verifier. The proposed price is 49 PLN for a license to use it in your internal tests. If you want the contents and purchase terms, ask in the comments; this is my own product, not an independent recommendation. There is no checkout link in this post. Please do not post customer data or payment information.

Top comments (0)