A clothing catalog is mostly photos, and the photos rarely come with good labels. A seller writes "nice outfit". A supplier sends a folder called new_final_2. Somebody then has to look at each picture and type "scarf".
This walkthrough turns that folder into a CSV: one row for every garment and accessory found, with the item type, a confidence score and a bounding box. It uses the API4AI Fashion API, about 80 lines of Python, and no ML setup.
At the end there's a short second script that does the part people skip: deciding which tags to trust and which photos a person should still look at.
One request
curl -X POST "https://api4ai.cloud/fashion/v2/results" \
-H "X-API-KEY: $API4AI_KEY" \
-F "image=@outfit.jpg"
The photo goes in as multipart form data: a file in image, or a public link in url. JPEG or PNG, under 16 MB.
What comes back
This is the example from the docs:
{
"results": [
{
"status": { "code": "ok", "message": "Success" },
"name": "scarf.jpg",
"md5": "a447c0aa2b2b89aa6ffb488317c0ab1e",
"width": 1024,
"height": 768,
"entities": [
{
"kind": "objects",
"name": "fashion",
"objects": [
{
"box": [0.2627, 0.0462, 0.6867, 0.6496],
"entities": [
{
"kind": "classes",
"name": "classes",
"classes": { "scarf": 0.6910 }
}
]
}
]
}
]
}
]
}
Every detected item is one entry in objects:
-
boxis[x, y, width, height]in normalized coordinates, 0 to 1 from the top-left corner. Multiply by the image width and height to get pixels. -
classesmaps a class name to a confidence between 0 and 1.
Look at the confidence in the docs' own example: 0.69. The scarf was found, but not with certainty. That number is what the second half of this post is about.
One more thing to be clear on: you get the item type and its position. Not color, not material, not size, not brand.
Three things that catch people
-
A photo the service can't read still returns HTTP 200.
status.codeis"failure"andstatus.messagesays why (wrong file type, corrupted image, a URL that can't be downloaded). If you only callraise_for_status(), those photos vanish from your catalog without a trace. - An outfit photo returns everything in it. A model wearing a shirt, a sweater and a belt gives you three items, and maybe only one is for sale. The boxes help: the item that fills most of the frame is usually the product.
-
Zero items is a valid answer. A photo with no detections comes back as
"ok"with an empty list. Log it, or you can't tell "nothing found" from "never processed".
The script
It walks the folder (subfolders included), sends each photo, and appends rows to a CSV as it goes.
"""Tag a folder of clothing photos with the API4AI Fashion API and write a CSV.
export API4AI_KEY=... # from portal.api4.ai
python3 tag_catalog.py photos/ tags.csv
One row per detected item (or one row per photo with nothing found, or failed).
Run it again after a crash: photos already tagged are skipped, failed ones retried.
"""
import csv
import os
import pathlib
import sys
import time
import requests
URL = "https://api4ai.cloud/fashion/v2/results"
KEY = os.environ["API4AI_KEY"]
PHOTOS = {".jpg", ".jpeg", ".png"}
src, out = pathlib.Path(sys.argv[1]), pathlib.Path(sys.argv[2])
def call_api(photo, tries=4):
for attempt in range(tries):
with photo.open("rb") as f:
r = requests.post(URL, headers={"X-API-KEY": KEY},
files={"image": f}, timeout=120)
if r.status_code in (429, 500, 502, 503, 504) and attempt < tries - 1:
time.sleep(2 ** attempt) # 1, 2, 4 s
continue
r.raise_for_status()
return r.json()["results"][0]
def items(result):
"""Yield (class name, confidence, box) for every detected item."""
for entity in result["entities"]:
if entity["kind"] != "objects":
continue
for obj in entity["objects"]:
classes = obj["entities"][0]["classes"]
name = max(classes, key=classes.get)
yield name, classes[name], obj["box"]
done = set()
if out.exists():
with out.open(newline="") as f:
done = {row["file"] for row in csv.DictReader(f) if row["status"] == "ok"}
photos = sorted(p for p in src.rglob("*") if p.suffix.lower() in PHOTOS)
counts = {"ok": 0, "skipped": 0, "failed": 0}
with out.open("a", newline="") as f:
writer = csv.writer(f)
if out.stat().st_size == 0:
writer.writerow(["file", "status", "item", "confidence", "x", "y", "w", "h"])
for photo in photos:
rel = str(photo.relative_to(src))
if rel in done:
counts["skipped"] += 1
continue
result = call_api(photo)
# A photo the service can't read comes back as HTTP 200 with "failure".
if result["status"]["code"] != "ok":
counts["failed"] += 1
writer.writerow([rel, "failed: " + result["status"]["message"]])
continue
counts["ok"] += 1
found = list(items(result))
if not found:
writer.writerow([rel, "ok", "", "", "", "", "", ""])
for name, confidence, box in found:
writer.writerow([rel, "ok", name, f"{confidence:.2f}",
*(f"{v:.3f}" for v in box)])
f.flush()
print(f"{len(found)} item(s) {rel}")
print(f"\n{len(photos)} photos: {counts['ok']} tagged, "
f"{counts['skipped']} already done, {counts['failed']} failed")
Run it:
pip install requests
export API4AI_KEY=...
python3 tag_catalog.py photos/ tags.csv
And the CSV looks like this:
file,status,item,confidence,x,y,w,h
look-01.jpg,ok,scarf,0.99,0.262,0.253,0.115,0.401
look-01.jpg,ok,belt,0.72,0.268,0.577,0.118,0.150
broken.jpg,failed: Can not load image.
flatlay-07.jpg,ok,,,,,,
What it takes care of:
- Resume. Photos that are already tagged are skipped, so after a crash or Ctrl-C you run the same command again. Failed photos are tried again; the old "failed" row stays in the file as a log.
- Rate limits and hiccups. A 429 or a 5xx waits 1, 2, then 4 seconds before giving up.
-
Failures you can see. Anything that comes back as
"failure"gets a row with the reason. - Empty photos. A photo with nothing found gets a row with an empty item.
- The count at the end. Tagged plus already-done plus failed should equal the number of photos.
It sends one photo at a time, which keeps the code short. For a large folder, wrap call_api in a ThreadPoolExecutor and lower the worker count if you start seeing 429s.
The part that matters: which tags to trust
A CSV of detections isn't a catalog yet. Between them sits one decision about confidence:
import csv
from collections import Counter
AUTO, REVIEW = 0.90, 0.60 # your thresholds, not ours
rows = list(csv.DictReader(open("tags.csv")))
items = [r for r in rows if r["status"] == "ok" and r["item"]]
auto = [r for r in items if float(r["confidence"]) >= AUTO]
review = [r for r in items if REVIEW <= float(r["confidence"]) < AUTO]
print(f"{len(auto)} tags to write automatically, {len(review)} for a person")
print("Catalog by item type:", Counter(r["item"] for r in auto).most_common())
empty = sorted({r["file"] for r in rows if r["status"] == "ok" and not r["item"]})
tagged = {r["file"] for r in rows if r["status"] == "ok"}
failed = sorted({r["file"] for r in rows if r["status"].startswith("failed")} - tagged)
print(f"{len(empty)} photos with nothing found, {len(failed)} failed")
Above the first threshold, write the tag. Between the two, show the photo to a person with the tag already suggested. Below, ignore it.
The 0.90 and 0.60 here are placeholders. Flat-lay shots on white behave differently from crowded street photos, so run a few hundred of your own images, look at what lands in each band, and move the numbers. Because the raw scores are in the CSV, changing a threshold later costs nothing: you don't call the API again.
Two more uses of the same file:
- Map names once. The API says "trousers"; your shop may say "pants". Keep a small dictionary in your code, not in the data.
-
Crop for visual search.
x * width,y * heightand so on give you a pixel box for each item, ready to cut out and index.
What it costs
On the API4AI developer portal the Fashion API is pay-as-you-go, with a list price of $300 per 1,000 requests and no subscription. One photo is one request.
That price is for small and occasional use. For a real catalog, ask for a quote: it depends on how many images you process, their resolution and how many items a typical photo contains, and at volume the discount is significant. Write to hello@api4.ai with a rough monthly number and a few sample photos.
Try it before you write any code
The Fashion API page has a demo that takes your own JPEG or PNG and draws the boxes with their scores. If the results look right on your kind of photo, a key on the portal takes a minute and needs no credit card. The full reference is in the docs, and the less technical version of this post, with the use cases, is on our blog.
How are you tagging product photos today: by hand, from the supplier feed, or with a model of your own?
Top comments (0)