DEV Community

Nishant Kumar
Nishant Kumar

Posted on

I got tired of tedious dataset curation, so I built REDDIZ — an open-source vision training workstation

Hacktoberfest: Maintainer Spotlight

Over the past few months, I've spent an unhealthy amount of time fine-tuning Flux and SDXL LoRAs and building small custom YOLO detectors. Every single time I start a new training run, the part that drives me crazy isn't setting learning rates or configuring optimizer schedules—it's the sheer manual grind of gathering and formatting the dataset.

If you train vision models, you probably know the drill all too well:

  1. You find a subreddit or niche photo board with great source material.
  2. You save images one by one or run an ad-hoc Python script that breaks halfway through due to Reddit's 429 rate limits or CORS blocks.
  3. You open a photo editor to manually crop every picture into consistent vertical portrait crops (9:16).
  4. You open a text editor to write out descriptive captions, camera tags, and trigger words for each image.
  5. If you're doing object detection, you boot up Label Studio, Roboflow, or CVAT, wait for them to load, and click around dozens of UI menus just to draw a few bounding boxes and export standard coordinate files.

I got tired of jumping across three separate tools and wrestling with paywalled export limits. I wanted a fast, standalone, keyboard-driven workstation where I could search subreddits, frame images at 9:16, run multi-modal AI models for auto-tagging, draw bounding boxes, and hit one button to get a clean .zip ready for Kohya_ss, OneTrainer, or Ultralytics YOLO.

That's why I built REDDIZ.


What REDDIZ Does (And How It Works)

REDDIZ is built around an authentic Swiss Modernist Brutalist layout using Archivo and Space Mono typography. Everything is laid out on a single, high-contrast screen with no nested menus or page reloads.

Here is a breakdown of what you can do inside the studio:

1. Ingestion: Multi-Subreddit Scraper & Local Uploader

You can harvest images directly from Reddit without needing a registered Reddit developer app or API client secret.

  • Type in any subreddit name or a combination separated by commas or spaces (for example: streetphotography, analog, EarthPorn).
  • Choose between sorting by NEW, HOT, or TOP.
  • Click FETCH MEDIA or press Enter. The built-in Node.js backend streams the posts through a lightweight proxy that strips out CORS restrictions, normalizes user agents to avoid 429 rate limits, and buffers full-resolution image URLs directly into the client.
  • If you already have your own photos, click UPLOAD MEDIA to batch drag-and-drop local .jpg/.png files or paste direct image URLs straight into the current session.

2. 9:16 Adaptive Centered Viewport & Resizable Panels

Diffusion models for character portraits, mobile wallpapers, and short-form video generation (like Wan 2.1 or HunyuanVideo) perform best when trained on consistent aspect ratios.

  • REDDIZ automatically centers incoming media into a normalized 9:16 crop.
  • The left column houses a vertical thumbnail filmstrip where you can quickly scrub through dozens of queued images.
  • Between the columns sits an interactive panel resizer. You can click and drag the divider horizontally to make the canvas larger on high-res monitors or expand the metadata side panel when working with long descriptive captions.

3. Multi-Modal AI Prompting & Auto-Captioning

Instead of forcing a single model or running a local Python server, REDDIZ connects directly to the provider of your choice using your own API key:

  • Google Gemini: Gemini 2.0 Flash, Gemini 2.0 Pro, Gemini 1.5 Pro
  • Anthropic Claude: Claude 3.5 Sonnet, Claude 3.5 Haiku
  • OpenAI: GPT-4o, GPT-4o Mini
  • Groq Cloud: Llama 3.2 11B & 90B Vision (incredibly fast, sub-second responses)
  • Hugging Face Hub: Serverless BLIP-2, ViT, and DETR endpoints

When you press + PROMPT (or use the shortcut Ctrl + Space), the vision engine analyzes the active image in real time. It looks at subject poses, facial features, clothing fabrics, environmental lighting, shadows, and camera perspective, then generates:

  1. A rich natural language caption for diffusion training.
  2. A list of granular token tags separated into individual chips.

Your API keys stay strictly inside your browser's localStorage. They are never sent to any intermediary server or database.


4. Bounding Box Annotation with Class Inheritance

For object detection datasets, you can draw 2D spatial bounding boxes directly over the canvas:

  • Built-in default classes: subject, foreground, background.
  • You can create custom classes on the fly with custom colors.
  • Auto-Inheritance: If you are annotating a video sequence or a photoshoot where the subject stays in roughly the same position across multiple frames, REDDIZ can inherit the bounding boxes from the previous image so you only have to nudge them instead of drawing from scratch.

5. Keyboard-First Ergonomics

The entire workflow is mapped to single-key shortcuts so you never have to move your hand back and forth between the keyboard and mouse:

  • [Enter] : Save current annotations and immediately load the next image in the queue.
  • [X] : Ignore/skip low-quality or irrelevant images.
  • [↑] / [↓] or [A] / [D] : Step backwards and forwards through the filmstrip.
  • [Ctrl + Space] : Trigger the active AI vision model.
  • [Esc] : Close any open modal dialog.

6. Universal 1-Click Export (.ZIP)

When you're done with a batch, click DONE & DOWNLOAD ZIP. REDDIZ bundles the entire dataset client-side:

  • Images: High-res, 9:16 cropped .jpg files numbered sequentially (0001.jpg, 0002.jpg, etc.).
  • Diffusion Text Pairs: Matching .txt files containing the prompt and token tags, ready to drop straight into Kohya_ss, OneTrainer, or a ComfyUI training pipeline.
  • YOLO Detection Labels: Matching .txt files with standard normalized coordinates:
  <class_id> <x_center> <y_center> <width> <height>
Enter fullscreen mode Exit fullscreen mode
  • COCO & JSONL Manifests: An annotations.json file adhering to COCO formatting and a dataset_manifest.jsonl file for programmatic pipelines.

Real-World Test: Annotating Classical Indian Artwork

To test REDDIZ end-to-end, I loaded a high-detail classical painting featuring two Indian women in traditional sarees standing beside a serene lake at sunset beneath massive cumulus clouds.

Here is the source image used for the test:

Test Art Source
Source image: Neoclassical landscape painting featuring two figures in traditional lehengas/sarees by the lake under sunset cumulus clouds.

Running the Annotation Inside REDDIZ

I dropped the image URL into REDDIZ, locked the 9:16 crop around the two women, and triggered Gemini 2.0 Flash with a single keystroke.

Here is the active studio screenshot captured during the test:

REDDIZ Annotation Studio in Action
Active annotation session: 9:16 cropped canvas, vertical filmstrip, extracted token chips (traditional_attire, indian_saree, cumulus_clouds, golden_hour, river_reflection), spatial classes, and live caption stream.

Results Generated by the Studio:

1. Generated Natural Language Caption:

"A tranquil neoclassical landscape painting capturing two Indian women in traditional draped sarees standing by a calm lake meadow at golden hour, beneath sweeping pastel-pink cumulus clouds and rolling distant hills."

2. Token Tags Extracted:
traditional_attire, indian_saree, cumulus_clouds, golden_hour, river_reflection, neoclassical_painting, scenic_landscape

3. Normalized YOLO Bounding Boxes:

  • subject (the two women): 0.4700 0.7900 0.1400 0.1600
  • background (clouds & hills): 0.1900 0.4300 0.7600 0.2800
  • foreground (flower meadow): 0.0000 0.9000 1.0000 0.1000

The entire process took less than 15 seconds from raw image ingestion to ready-to-train export files.


Technical Stack & Philosophy

  • Zero-Bundle Frontend: Written in clean, vanilla HTML5, CSS3, and ES6+ JavaScript. No React, no Vue, no bloated virtual DOM. The page boots instantly in under 50ms.
  • Typography & Theme: Built using Google Fonts Archivo for headings and editorial copy, paired with Space Mono for technical coordinates and code.
  • Privacy First: Everything runs inside your browser. No cookies, no tracking scripts, no third-party telemetry. Your vision API keys stay on your machine.
  • Deployment: Hosted on Vercel with a lightweight Node.js streaming proxy for handling external media headers and Reddit API pagination.

Try It Out

REDDIZ is completely free and open source:

If you build or fine-tune vision models, give it a run with your favorite subreddits or image folders. Pull requests and feedback are always welcome!

Top comments (0)