The Ledger That Stops You Double-Posting Your Stock Uploads
I still remember the sinking feeling when I opened my Adobe Stock contributor dashboard and saw the same 47 vector files sitting in "pending review" twice. I had uploaded them once on a Monday, then my batch script crashed halfway through. I assumed it had failed completely, so I ran the whole thing again on Tuesday. Forty-seven duplicate uploads. My approval rate tanked, I got a warning email, and I spent an entire weekend manually deleting submissions.
That was the day I built upload_tracker.py. Not because tracking sounded like a nice idea, but because I never wanted to see that duplicate list again. It is one of four standalone scripts in what I now call the Pipeline Starter Kit, and it is the one that saved my account.
The real rule nobody tells you: a rerun is not idempotent
Most upload tooling assumes you will run it once. Batch scripts that build metadata CSVs, zip assets, and POST them to a platform's upload endpoint usually have no memory. If the process dies at file 30 of 100, you have no clean way to resume. You either manually figure out which 29 made it through, or you rerun the lot and eat the duplicates.
Adobe, Vecteezy, and Dreamstime all treat a duplicate submission as a strike against you, even if it is accidental. The fix is not a clever retry loop. It is a ledger: a file that records what has already been posted, and a pending command that only ever hands you the files that are not in it.
Here is the whole workflow:
python upload_tracker.py pending items.txt # shows what is NOT done yet
python upload_tracker.py done "asset-001.jpg"
python upload_tracker.py status
That is it. You pipe pending into whatever uploads your files, and you call done on each one that succeeds. If the process crashes, you rerun pending and it picks up exactly where you left off. No duplicates, no manual archaeology.
Why the CSV headers matter more than your art
The second thing that gets accounts flagged is a malformed upload CSV. All three major channels reject a wrong header instantly, and each one wants a different shape: Adobe takes 5 columns, Vecteezy takes 4, Dreamstime takes 15. Get one wrong and the whole file bounces.
make_metadata.py exists purely to write the header each platform demands, and nothing else:
python make_metadata.py ./my-assets --platform adobe --out metadata.csv --ext .jpg
python make_metadata.py ./my-assets --platform vecteezy --out vecteezy.csv --ext .jpg
python make_metadata.py ./my-assets --platform dreamstime --out dreamstime.csv --ext .jpg
Without --titles titles.json, it title-cases the filename and uses that. That is fine for a first pass. Replace it before you submit, because title-cased filenames do not sell.
Two dependencies, and you can skip both
The kit has four scripts and only two third-party packages total: Pillow for the covers in pack_product.py, and reportlab for md2pdf.py. If you only use make_metadata.py and upload_tracker.py, you need nothing but Python itself. No pip install, no virtualenv, no cloud account.
That matters when you are running on cheap hardware. This entire pipeline ran on a single-core, 2 GB server with no GPU, and it packaged 39 products producing 856 MB of output. It works because of a few hard habits:
- One process at a time. Never fan out image jobs on a 1-core box.
- Call
gc.collect()after each asset to free memory between items. - Measure peak RSS, not average, using
resource.getrusage(RUSAGE_SELF).ru_maxrss. - Run heavy render steps as one file per subprocess so memory returns to the OS.
- Render previews at 60 dpi when you are inspecting, not 300.
These are not clever tricks. They are what you learn after watching a 2 GB box swap itself to death because you thought parallel uploads would be faster.
Packaging is where most products give up
A zip with a real cover sells. A bare file does not. pack_product.py does both in one command:
python pack_product.py my-product "My Product" "A short subtitle" 9
It writes the zip, a 1280x720 cover, and a 600x600 thumbnail into a _system folder. Point it elsewhere with --dir or the PACK_ROOT environment variable; no need to edit the script. Meanwhile md2pdf.py turns your content into something that looks like a product instead of a raw markdown dump:
python md2pdf.py content.md product.pdf "My Product Title"
The folder layout that keeps a catalogue from becoming a junk drawer
project/
assets/ raw output (one file per asset)
output/ ready-to-submit files + metadata.csv
packs/ finished product zips + covers
research/ your notes and sources
scripts/ these tools
Keep research/ forever. Your scrape logs and notes are raw material for future products. I deleted mine once to "clean up" and I will not make that mistake again. The 300+ pages of research behind the rules these scripts enforce only exist because I stopped deleting things.
The honest part
None of this is a magic revenue machine. It is four scripts that solve the four problems that actually stall solo makers: metadata headers, double-posting, turning docs into products, and packaging. The kit is the working version of everything above, built from roughly 19,000 lines of pipeline code in my parent project and validated on that same 1 vCPU, 2 GB box.
If you have been uploading by hand or running a script that you are afraid to rerun, the ledger alone is worth the download.
Building digital products? Free guides & kits for digital product sellers.
Top comments (0)