You get a .parquet file and need to check what's in it. No Python, no Spark, ideally without uploading it anywhere. Here is a quick demo with the free DataToolbox Parquet Viewer, using a test file built to break things.
1. Open it
Drop the file on the viewer (or click Try sample file). The footer is read first, so you immediately see rows, columns and row groups. Pages read only the row groups they need.
2. Check the values that usually go wrong
| Column type | Stored | Shown |
|---|---|---|
| INT64 | 9007199254740993 |
9007199254740993 (not ā¦992) |
| DECIMAL(38,10) | integer + scale | 1234567890123456789012345678.0123456789 |
| TIMESTAMP(NANOS, UTC) | INT64 ns | 2026-03-08T07:30:00.123456789Z |
| TIMESTAMP(MICROS, no timezone) | INT64 µs |
2026-03-08T02:30:00.123456 (no Z) |
| INT96 (Spark's default) | 12 bytes |
2026-03-08T07:30:00.123456000, labelled INT96, no zone |
NULL, empty string and a missing struct field are displayed differently.
3. Export
Parquet to CSV and Parquet to JSON let you choose the current page, the filtered rows or all rows. CSV has a delimiter option, a NULL token (empty, NULL, \N) and nested columns as JSON text or left out; empty strings are written as "". JSON writes decimals, and integers beyond ±2^53ā1, as strings.
Limits (one machine)
On an Apple M4 / 16 GB / Chromium: 5,000,000 rows opened in about 0.1 s and exported to a 338 MB CSV in about 9 s. Reads above 400 MB and exports above 250 MB of decoded data are refused, because larger ones froze or crashed the tab in testing. Your device may hit limits sooner.
Full walkthrough, including the DuckDB equivalents: How to open and inspect a Parquet file without Python or Spark.
Top comments (0)