DEV Community

Maya Chen
Maya Chen

Posted on

Checking missing dates in a legal metadata CSV with Python

A small Python audit that separates missing source dates from malformed values, preserves judge-page URLs and reports what a dated metadata snapshot can actually show.

A court-rules index is a list of pages with dates attached, and a date is the field most likely to be empty. The Court Rules Judge Index snapshot for September 28, 2026 has 1,233 rows across 39 courts, and 486 of them carry no date at all. Before quoting any count from a file like this, run a quick audit of the dates.

The snapshot

The CSV has nine columns. The one this post cares about is rules_last_changed, an ISO date such as 2024-09-01, or an empty cell when the judge page shows no date.

column meaning
judge_name the judge or commissioner on the page
judge_status Judge, Senior Judge, Magistrate Judge and similar
court_name the court the page belongs to
court_slug the court's short id, for example sdny
court_type federal or state
rule_topics the rule topics the page covers, separated by semicolons
rule_count how many rule records the page lists
rules_last_changed the date the page displays, or empty
url the judge page on courtrules.app

The audit

The script uses only the standard library. It never writes today's date into a missing cell, and every line of the report repeats the judge-page URL so a finding can be checked against the page.

#!/usr/bin/env python3
"""Audit a judge index CSV for missing, malformed and present source dates.

Usage: python3 date-audit.py [judges.csv]

Reads the CSV named on the command line (default: sample-judges.csv) and checks
rules_last_changed against ISO YYYY-MM-DD dates. An empty value is missing, an
unparseable value is malformed, everything else is present. Today's date is
never substituted for a missing value, and every report line keeps the row's
judge-page URL.

Standard library only.
"""
import csv
import datetime
import sys

COLUMNS = ("judge_name", "court_slug", "url", "rules_last_changed")


def is_iso_date(value):
    try:
        datetime.date.fromisoformat(value)
        return True
    except ValueError:
        return False


def audit(path):
    missing = []
    malformed = []
    present = []
    with open(path, newline="", encoding="utf-8") as handle:
        reader = csv.DictReader(handle)
        fields = reader.fieldnames or []
        absent = [name for name in COLUMNS if name not in fields]
        if absent:
            print("cannot audit, the input has no " + ", ".join(absent))
            return 2
        for row in reader:
            value = (row.get("rules_last_changed") or "").strip()
            record = (row.get("judge_name", ""), row.get("court_slug", ""), row.get("url", ""), value)
            if value == "":
                missing.append(record)
            elif is_iso_date(value):
                present.append(record)
            else:
                malformed.append(record)
    total = len(missing) + len(malformed) + len(present)
    print("rows read = " + str(total))
    print("present = " + str(len(present)) + ", missing = " + str(len(missing)) + ", malformed = " + str(len(malformed)))
    print("")
    for label, records in (("missing date", missing), ("malformed date", malformed), ("present date", present)):
        print(label + " (" + str(len(records)) + ")")
        for name, court, url, value in records:
            shown = value if value else "(empty)"
            print("  " + court + " | " + name + " | " + shown + " | " + url)
        print("")
    return 1 if missing or malformed else 0


if __name__ == "__main__":
    target = sys.argv[1] if len(sys.argv) > 1 else "sample-judges.csv"
    sys.exit(audit(target))
Enter fullscreen mode Exit fullscreen mode

Run it

python3 date-audit.py judges.csv
Enter fullscreen mode Exit fullscreen mode

Against the real snapshot the summary is one line, and the per-row blocks list every missing and malformed value.

rows read = 6
present = 3, missing = 1, malformed = 2

missing date (1)
  ex-district | Example Judge Missing Date | (empty) | https://example.invalid/courts/ex-district/example-judge-missing-date

malformed date (2)
  ex-district | Example Judge Month Name | March 2024 | https://example.invalid/courts/ex-district/example-judge-month-name
  ex-district | Example Judge Impossible Date | 2024-02-30 | https://example.invalid/courts/ex-district/example-judge-impossible-date

present date (3)
  ex-district | Example Judge, With Comma | 2024-09-01 | https://example.invalid/courts/ex-district/example-judge-with-comma
  ex-district | Example Judge Padded Date | 2024-09-01 | https://example.invalid/courts/ex-district/example-judge-padded-date
  ex-district | Example Judge Old Date | 2019-02-01 | https://example.invalid/courts/ex-district/example-judge-old-date
Enter fullscreen mode Exit fullscreen mode

The output above comes from a synthetic fixture beside the script. It carries the cases that matter. A quoted name with a comma parses as one field. A padded value strips to a clean date. March 2024 and 2024-02-30 both fail the ISO check even though a human reader might accept one of them. The real file, by contrast, reports 747 present dates and 486 missing ones, and no malformed values.

What the audit cannot show

  • A missing date stays unknown. The script leaves the cell empty and never guesses.
  • A present date describes what the page displayed when the snapshot was built. Legal deadlines come from the court's current rules.
  • The count measures rows in one file. It says nothing about whether a court requirement applies to a case.
  • The audit checks the file only. The court documents behind the pages need a separate source check.

Source and license

The snapshot is published at Alison J. Nathan's source page and each row links its page. Nathan's entry traces back to the court's own criminal rules PDF, which is the kind of document the pages summarize. CC BY 4.0 covers judges.csv, judges.json and the dataset README only. Credit Court Rules and keep every row's source URL.

The script is short enough to adapt: point it at a newer snapshot, or swap the column name for any other dated field, and it reports the same three buckets.

Top comments (1)

Some comments may only be visible to logged-in visitors. Sign in to view all comments.