Imagine reviewing a configuration file for a nightly import job.
You expect the job to stop after processing 1,000 rows. But the application keeps going until 10,000.
The file looks reasonable at first glance:
{
"task": "nightly-import",
"limits": {
"max_rows": 1000,
"max_rows": 10000
}
}
There are no missing commas or broken brackets. Your JSON parser doesn't throw an error.
Look again at max_rows.
The key appears twice, with two different values.
That small detail changes what the application actually receives.
What JavaScript does with this input
Save this as check.js and run it with Node.js:
const source = `{
"task": "nightly-import",
"limits": {
"max_rows": 1000,
"max_rows": 10000
}
}`;
const config = JSON.parse(source);
console.log(config.limits.max_rows);
// 10000
console.log(Object.keys(config.limits));
// [ 'max_rows' ]
console.log(JSON.stringify(config, null, 2));
The final output contains only one max_rows property:
{
"task": "nightly-import",
"limits": {
"max_rows": 10000
}
}
The earlier value didn't survive parsing.
This is the part that makes the bug inconvenient: if you log only the parsed object, the evidence is gone.
The log shows a normal object with one property. Nothing tells you that the original payload contained two competing values.
If the raw document came from a generated configuration, a webhook, or a file assembled from multiple sources, the mistake may have happened several steps earlier.
Is duplicate-key JSON actually invalid?
RFC 8259, Section 4, says that member names within a JSON object SHOULD be unique.
That's an important distinction from MUST be unique.
The specification explains that implementations can handle duplicate names differently. Some keep the last occurrence; others reject the document or preserve multiple pairs.
So the safer engineering assumption is not "every parser will reject this."
It's "a document with duplicate object names may not be interpreted consistently."
For configuration files and API contracts, I prefer a stronger rule: reject duplicate names before the application uses the parsed values.
Why checking the object afterward doesn't work
You might try validating config after calling JSON.parse().
But by then, JavaScript has already reduced the duplicate entries to a single property.
A JSON.parse() reviver can't recover the overwritten member either. It receives values from the resulting parse, not an ordered record of every original occurrence.
The same limitation applies when a JSON Schema validator is given an already-parsed object. It can check the surviving value's type, range, and required fields, but it cannot reconstruct a duplicate member that was discarded before validation.
This is why duplicate-key detection belongs at the parsing boundary.
A small Python parser that rejects duplicates
Python's standard json module provides object_pairs_hook.
Instead of immediately reducing every JSON object to a dictionary, this hook receives its ordered member pairs.
That gives us a chance to reject duplicates.
import json
def reject_duplicates(pairs):
result = {}
for key, value in pairs:
if key in result:
raise ValueError(
f"Duplicate JSON key: {key!r}"
)
result[key] = value
return result
def load_unique_json(source):
return json.loads(
source,
object_pairs_hook=reject_duplicates,
)
Now test a duplicate in a nested object:
source = '''
{
"task": "nightly-import",
"limits": {
"max_rows": 1000,
"max_rows": 10000
}
}
'''
load_unique_json(source)
Instead of choosing one value, it raises:
ValueError: Duplicate JSON key: 'max_rows'
The hook also runs for nested objects, including objects inside arrays.
We don't need to scan the raw file with regular expressions or write an entire JSON parser.
Three regression cases I would keep
A duplicate detector needs to reject repeated names within the same object, not across unrelated objects.
This should pass:
valid = '''
{
"primary": {"id": 1},
"backup": {"id": 2}
}
'''
assert load_unique_json(valid) == {
"primary": {"id": 1},
"backup": {"id": 2},
}
These should fail:
invalid_cases = [
# Duplicate in a top-level object
'{"mode": "safe", "mode": "fast"}',
# Duplicate inside an array element
'{"items": [{"id": 1, "id": 2}]}',
# Same decoded key, different source spelling
r'{"name": "first", "\u006eame": "second"}',
]
for source in invalid_cases:
try:
load_unique_json(source)
except ValueError as error:
print("Rejected:", error)
else:
raise AssertionError(
"Duplicate key was not rejected"
)
The third case is particularly easy to overlook.
In JSON, \u006e represents the letter n.
So these two names:
"name"
"\u006eame"
decode to the same string.
A simple search for repeated quoted text can miss the duplicate. The parser-based approach checks the decoded names instead.
This test is about exact decoded-key equality. It does not attempt to detect visually similar Unicode characters or perform Unicode normalization.
One limitation worth knowing
Our Python function detects duplicate keys, but its error only reports the member name.
It does not tell you the exact path, line, or column.
For a large configuration containing many objects with fields named id or timeout, that may not be enough to locate the source.
A production-grade diagnostic tool should ideally report something like:
Duplicate member:
path: $.limits.max_rows
first occurrence: line 4
second occurrence: line 5
That requires retaining more source-location information than this short example does.
For a small configuration or a CI test, though, rejecting the document is already much better than silently choosing a value.
Put the check where data enters your application
For a configuration-loading pipeline, my preferred order is:
- Read the original JSON text.
- Reject duplicate object-member names.
- Parse the accepted document.
- Validate its schema and application-specific constraints.
- Only then use it to configure the application.
In the Python example, steps 2 and 3 happen together.
The important point is to avoid discarding evidence before you have checked the rule you care about.
A browser-based syntax checker is still useful for malformed JSON. For example, I maintain a JSON Validator on DataToolForge for ordinary syntax and structure checks. That is a different task from the duplicate-key rejection implemented in this article.
A successful parse is useful evidence that the syntax was accepted.
It is not evidence that every member of the original document survived.
References
The configuration example is synthetic. The JavaScript and Python examples were tested with Node.js 22.16.0 and Python 3.13.5.
Disclosure: DataToolForge is my project. AI assistance was used to draft, research, and test the examples in this article.
Top comments (0)