A ZIP archive can contain more than one entry with the same name. A name alone then cannot identify which entry you meant. This read-only checker reports repeated entry names and their positions without extracting or changing anything.
1. Make practice archives
A ZIP archive stores entries and a directory describing them. Two entries can use the same name even when their contents differ. Save this as MakeExamples.py in a practice folder:
import warnings as Warnings
import zipfile as ZipFile
from pathlib import Path
Names = [Path("Repeated.zip"), Path("Distinct.zip")]
if any(Name.exists() or Name.is_symlink() for Name in Names):
raise SystemExit("Use a fresh practice folder without Repeated.zip or Distinct.zip.")
with Warnings.catch_warnings():
Warnings.simplefilter("ignore", UserWarning)
with ZipFile.ZipFile(Names[0], "x") as Archive:
Archive.writestr("Note.txt", b"first")
Archive.writestr("Note.txt", b"second")
with ZipFile.ZipFile(Names[1], "x") as Archive:
Archive.writestr("Note.txt", b"first")
Archive.writestr("note.txt", b"second")
Run:
python3 MakeExamples.py
The setup stops if either archive name already exists, including as a symbolic link. Mode "x" creates a new archive rather than replacing an existing file. The warning filter is only around the intentional repeated-name example: Python warns when that second Note.txt is written. Use these small practice archives rather than real files you need.
2. Save the name checker
Save this as CheckZipNames.py beside the archives:
import json as Json
import sys as Sys
import zipfile as ZipFile
from pathlib import Path
def CheckZipNames(FileName):
with Path(FileName).open("rb") as Input:
Input.seek(0, 2)
if Input.tell() > 1048576:
raise ValueError("archive exceeds 1 MiB teaching limit")
Input.seek(0)
with ZipFile.ZipFile(Input) as Archive:
Entries = Archive.infolist()
if len(Entries) > 1000:
raise ValueError("more than 1000 entries")
Groups = {}
for Number, Entry in enumerate(Entries, start=1):
Groups.setdefault(Entry.filename, []).append(Number)
Repeated = [{"Name": Name, "Entries": Numbers}
for Name, Numbers in Groups.items() if len(Numbers) > 1]
return {"EntryCount": len(Entries), "RepeatedNames": Repeated}
def Main():
if len(Sys.argv) != 2:
print("Usage: python3 CheckZipNames.py input.zip", file=Sys.stderr)
return 2
try:
Report = CheckZipNames(Sys.argv[1])
except (OSError, ValueError, ZipFile.BadZipFile) as Problem:
print(f"Name check stopped: {Problem}", file=Sys.stderr)
return 2
print(Json.dumps(Report, ensure_ascii=True, indent=2))
return 1 if Report["RepeatedNames"] else 0
if __name__ == "__main__":
Sys.exit(Main())
The archive limit is 1 MiB (1,048,576 bytes), measured on the opened file before the ZIP directory is parsed. This is the archive's stored size, not its total uncompressed size. Use a saved regular file that will not change while it is read.
infolist() returns an entry object for each directory entry, in archive order. The program groups each exact filename and records one-based entry positions. A dictionary keyed only by name would discard this distinction if it replaced earlier entries.
The 1,000-entry limit is checked after the library has loaded the ZIP directory. It limits the accepted report, not the work needed to parse that directory. These teaching limits are not a security sandbox or a guarantee about handling hostile archives.
3. Find the repeated name
Run:
python3 CheckZipNames.py Repeated.zip
The output is:
{
"EntryCount": 2,
"RepeatedNames": [
{
"Name": "Note.txt",
"Entries": [
1,
2
]
}
]
}
The exit code is 1. Both entries have the exact name Note.txt; their positions distinguish them even though the name does not. The checker does not choose one or compare their contents. The same report would appear if those contents matched.
JSON is a text format for named values and lists. The exit code is 0 when no exact repeated names are found, 1 when they are found and 2 when usage, reading, parsing or a teaching limit stops the check. No partial report is printed after a caught error.
4. Keep exact matching separate from extraction rules
Run:
python3 CheckZipNames.py Distinct.zip
The output is:
{
"EntryCount": 2,
"RepeatedNames": []
}
The exit code is 0. Note.txt and note.txt are different strings. That does not prove they would stay separate on a case-insensitive destination filesystem. This checker does not fold case, remove spaces or resolve path spellings. Directory entries are included too, so two exact directory names would be reported as a repeat.
The checker never reads entry contents, checks their CRC values, tests passwords or extracts files. A valid ZIP directory with no repeated names is not a claim that every entry is readable, the archive is complete or its paths are safe to extract. Do not extract an unfamiliar archive merely because this check returns 0.
5. Keep the report separate from the input
To save a report, use a new filename:
python3 CheckZipNames.py Repeated.zip > RepeatedNames.json
This creates or replaces RepeatedNames.json; keep that output name free. Never redirect output over the input archive, because the shell can empty it before Python opens it. Entry names can reveal personal or work information, so keep reports private when they describe real archives.
Top comments (0)