DEV Community

Arya Hegiste
Arya Hegiste

Posted on

Import CSV Data into DynamoDB with Tables: Validate, Fix, and Retry Invalid Rows

Importing CSV data into DynamoDB involves more than moving rows from a file into a table. Before writing anything, it is important to understand how source columns map to DynamoDB attributes, how types are inferred, and what happens when a record is missing a required key.

In this walkthrough, I use Tables by Serverless Creed to import CSV data into DynamoDB and examine how the workflow handles valid records, type inference, and rows with missing partition or sort keys.

We’ll specifically look at:

  • Importing a large CSV fixture
  • Reviewing type inference
  • Testing records with missing DynamoDB keys
  • Confirming invalid input is blocked before import
  • Correcting the input and retrying
  • Verifying the recovered records in DynamoDB The goal is not simply to see an “import completed” message. We want to verify what actually reached the table and understand how to recover when validation fails.

1. Prepare the import fixture
For this exercise, I used the disposable DynamoDB table:

TC-IMP-002-Test
Enter fullscreen mode Exit fullscreen mode

The table uses a composite key consisting of a partition key and sort key.

The main CSV fixture contained 1,000 source rows designed to exercise several import behaviors:

  • Valid records
  • Duplicate-key records
  • Records missing required keys
  • A value designed to inspect type inference

After the initial import, the DynamoDB table contained:

961 items
Enter fullscreen mode Exit fullscreen mode

This stored-item count should not automatically be interpreted as “961 accepted and 39 rejected.”

Duplicate DynamoDB keys can overwrite an existing item rather than increase the number of unique items in the table. Also, because the original import completion dialog was not retained, I do not have direct evidence for the exact accepted, rejected, and duplicate counters from that run.

For that reason, I use the resulting table state as an observation rather than reconstructing import-summary numbers that were not captured.

Figure 1 — DynamoDB table after the initial CSV import


2. Inspect type inference instead of assuming a validation error
One record in the fixture was intentionally given a value that looked incompatible with a numeric field:

PK: INVALID#001
Value: not-a-number
Enter fullscreen mode Exit fullscreen mode

A reasonable assumption might be that the importer would reject the record because not-a-number is not numeric.

That is not what happened in this test.

After the import, I inspected INVALID#001 in Tables. The value had been imported as a DynamoDB String.

In other words, the importer did not treat this value as a malformed number. Its type inference allowed the value to be represented as text.

Figure 2 — not-a-number imported as a String through type inference

This is an important distinction when investigating an import.

A value that appears invalid according to an expected application schema is not necessarily invalid according to DynamoDB's data model or the importer's inferred type.

Required key violations, however, are different.


3. Create a controlled missing-key test
To test a clear validation failure, I used a smaller CSV containing three rows:

PK,SK,Name,Value
TEST#001,PROFILE,Valid Control,1
,PROFILE,Missing PK,2
TEST#003,,Missing SK,3
Enter fullscreen mode Exit fullscreen mode

The first record contains both required keys.

The second record is missing the partition key (PK).

The third record is missing the sort key (SK).

This gives us one valid control record and two deliberately invalid records.

When this file was loaded into the import workflow, the preview made the missing key values visible before the import was allowed to proceed.

Figure 3 — Import preview containing missing PK and SK values


4. Block invalid rows before writing
Tables identified the key problem during preview validation.

The interface reported:

2 preview rows are missing a required key value. Fix the source data or mapping before importing.

The Start import action was disabled.

This is an important safety behavior because DynamoDB requires the complete primary key for an item. Rather than beginning the import and discovering the key problem after writes had started, the workflow blocked this controlled import during validation.

The screen also showed the target table and import configuration, making it possible to confirm where the operation would write before proceeding.

Figure 4 — Import blocked because two rows are missing required key values

At this point, the correct recovery was not to bypass the validation. The source data or mapping needed to be corrected.


5. Correct the source and retry
I corrected the three-row CSV so that every record had both required key values:

PK,SK,Name,Value
TEST#001,PROFILE,Valid Control,1
TEST#002,PROFILE,Fixed Missing PK,2
TEST#003,PROFILE,Fixed Missing SK,3
Enter fullscreen mode Exit fullscreen mode

After loading the corrected input, the preview changed to:

Preview validation passed. The streaming import is ready to run.

The previously blocked import could now proceed.

Figure 5 — Corrected input passes validation and is ready for retry

This creates a clear recovery sequence:

Missing required keys
        ↓
Preview validation fails
        ↓
Import blocked
        ↓
Correct source/mapping
        ↓
Preview validation passes
        ↓
Retry import
Enter fullscreen mode Exit fullscreen mode

No special workaround was needed. The recovery consisted of correcting the data so it satisfied the target table's key requirements and then rerunning the normal import workflow.


6. Verify the recovered data in DynamoDB
An import completion message is useful, but I also wanted to verify that the corrected data actually existed in DynamoDB.

After the retry, I read back:

PK: TEST#002
SK: PROFILE
Enter fullscreen mode Exit fullscreen mode

The resulting item contained:

Name: Fixed Missing PK
Value: 2
Enter fullscreen mode Exit fullscreen mode

The table count had also increased from:

961
Enter fullscreen mode Exit fullscreen mode

to:

964
Enter fullscreen mode Exit fullscreen mode

which is consistent with the three corrected test records being added as unique items.


Figure 6 — Corrected TEST#002 item read back after the successful retry

This final read-back is useful because it verifies the resulting table state rather than relying solely on the import UI.


Type inference and key validation solve different problems
One of the more useful lessons from this exercise is the difference between type inference and required-key validation.

Consider:

Value = not-a-number
Enter fullscreen mode Exit fullscreen mode

The importer could represent that value as a string, so it was not necessarily invalid for DynamoDB.

But this:

PK = <missing>
Enter fullscreen mode Exit fullscreen mode

cannot produce a valid item for a table whose primary key requires PK.

That is why the second case produced a blocking validation error while the first did not.

When investigating rejected or unexpected CSV imports, it helps to separate these questions:
1. Can the source value be represented as a DynamoDB type?
2. Does the resulting item satisfy the target table's required key structure?

They are related to import validation, but they are not the same check.


Why read-back verification matters
Import tools can tell us whether an operation started, completed, or encountered validation errors.

But when correctness matters, it is useful to verify the resulting DynamoDB state independently.

For this exercise, the strongest recovery evidence was not simply that the corrected preview said:

Preview validation passed.

It was that the corrected item could subsequently be read from the table with the expected key and values.

A useful validation workflow is therefore:

CSV source
   ↓
Preview and mapping
   ↓
Validation
   ↓
Import
   ↓
DynamoDB read-back
Enter fullscreen mode Exit fullscreen mode

If validation fails, correct the source or mapping and repeat the same path rather than assuming the failed records were written.


Conclusion
CSV imports into DynamoDB become easier to troubleshoot when validation is treated as part of the workflow rather than just a gate before the import.

Using Tables, I was able to inspect inferred values, reproduce a controlled missing-key failure, see the import blocked before proceeding, correct the affected records, retry the operation, and verify the recovered data in DynamoDB.

The test also demonstrated why apparent data-type problems and required-key problems should be investigated separately. A value such as not-a-number may still be valid as a DynamoDB string, while a missing required partition or sort key prevents the item from satisfying the table's key schema.

Most importantly, don't stop verification at “Import completed.”

For repeatable CSV-to-DynamoDB workflows, a stronger process is:
preview → validate → import → read back.

That final read-back provides evidence of what actually reached DynamoDB.


Give it a try on https://tables.serverlesscreed.com/

Top comments (2)

Collapse
 
launchgatecheck profile image
Launch Gate •

Does "preview validation passed" cover the entire file, or just the displayed preview rows? I'd add a fixture with a missing SK well beyond the preview window, then check whether any earlier valid rows were written before the error was found. That makes the all-or-nothing versus partial-import contract visible.

For the duplicate-key case, I'd use two rows with the same PK/SK but different values and read back the final value, rather than use the table count to infer which row won. I haven't used Tables; these are follow-up fixtures suggested by the distinction you make between summary counters and stored state.

Collapse
 
palak_jain_427d9b8870a520 profile image
Palak Jain •

Insightful!!!