I Built a Metadata CSV Generator So I Stop Getting Rejected
The rejection email was three words long: "Wrong CSV format."
I had spent four hours uploading assets, and Adobe Stock bounced the entire batch at the validation step. Not a single file was reviewed. The problem was not the art. The problem was that my header row said Title,Keywords,Category when Adobe wanted Filename,Title,Keywords,Category,Releases. One extra column and the whole upload is dead on arrival.
That was the moment I stopped treating metadata as an afterthought. This post is about the four scripts I built to make sure it never happens again, and the platform rules that forced me to write them.
The CSV headers are not optional
Every stock channel enforces an exact column layout, and every one of them rejects a file with the wrong shape instantly. The formats I had to satisfy:
- Adobe Stock: 5 columns
- Vecteezy: 4 columns
- Dreamstime: 15 columns
Fifteen columns for Dreamstime. If you are hand-editing a spreadsheet for that, you will get it wrong eventually, and the failure mode is silent until the upload fires. So make_metadata.py writes the header the platform demands and nothing else. You point it at a folder, tell it which platform, and it emits the CSV. It also handles the boring case: if you do not pass a titles file, it title-cases the filename and uses that. That is fine for a first pass. Replace it before you submit, because a filename is not a title.
The other half is not posting the same batch twice. The fastest way to get an account flagged is to run your upload job again because you were not sure it went through. upload_tracker.py is a ledger. It remembers what is already done, so a re-run never double-posts. You ask it for pending and it shows you exactly what is left, not what you have already shipped.
Packaging is where products become products
A zip with a real cover sells. A bare file does not. pack_product.py bundles it: it zips the files and renders a cover at 1280x720 plus a thumbnail at 600x600. Pillow does the images. That is the whole dependency for this step.
The fourth script, md2pdf.py, turns a Markdown file into a styled PDF with headings, tables, lists and links intact. If you already have the content, this is the step that makes it look like something a person would pay for. reportlab handles the output.
Two third-party packages total across the kit: Pillow and reportlab. The metadata and tracker scripts use the standard library alone. No install needed for those two at all.
It all runs on a 2 GB server
This is the part people do not believe. The full pipeline ran on 1 vCPU, 2 GB RAM, no GPU. Not a demo. It packaged 39 finished products and produced 856 MB of output on that box.
The habits that make that possible:
- One process at a time. Never fan out image jobs in parallel on a single core. It does not go faster, it goes out of memory.
-
Free memory between items. Call
gc.collect()after each asset. -
Measure peak RSS, not average. Averages hide the spike that killed the job.
resource.getrusage(RUSAGE_SELF).ru_maxrsstells you the real ceiling. - One file per subprocess for heavy render steps, so memory is returned to the OS instead of held.
- Render at the resolution you need. Inspecting a preview at 60 dpi beats rendering at 300 and waiting.
For context on how much rule-enforcement is behind these four small tools: the parent project is roughly 19,000 lines of pipeline code, built on 300+ pages of research into what each platform actually accepts.
Keep the research folder
One structure decision matters more than the rest:
project/
assets/ raw output, one file per asset
output/ ready-to-submit files + metadata.csv
packs/ finished zips + covers
research/ your notes and sources
scripts/ these tools
Keep research/ forever. Your scrape logs and notes are the raw material for future products. Deleting them to "clean up" is a mistake this project made once and will not repeat, because the next product idea usually came out of a note that looked useless at the time.
What this actually is
I packaged the working scripts. Not a tutorial that ends at "here is how you would do it." These are the four tools that did the job, no framework and no cloud account:
-
make_metadata.pyfor the Adobe, Vecteezy and Dreamstime CSVs -
upload_tracker.pyfor the ledger that stops double-posting -
md2pdf.pyfor Markdown to styled PDF -
pack_product.pyfor the zip, cover and thumbnail
If you have been rejected on metadata and blamed your art, this is the fix.
Building digital products? Free guides & kits for digital product sellers.
Top comments (0)