DEV Community

Cover image for CPU probes for AI tattoo generators — what we measured
Lena Hart
Lena Hart

Posted on

CPU probes for AI tattoo generators — what we measured

By Lena Hart

tattooprobes resamples a design to a print size on CPU, measures lines and gaps in millimetres, and can blur the raster before it scores the skeleton. A PNG from an AI tattoo generator goes in as an array.

Why the probes stayed on CPU

probes.py and simulate.py import NumPy, OpenCV, and scikit-image. The locked environment is still a CPU torch build, because reward scoring and the CLIP labeler share the venv. requirements.lock.txt pins torch==2.14.1+cpu, open_clip_torch==3.3.0, transformers==4.46.3, and image-reward==1.5. The geometry pass ran without a GPU, without a paid scoring API, and without new annotators. Seed 0 is the simulator default, torch.manual_seed(0) in the classifier, and SEED = 0 in the fusion script.

Two functions

Excerpt from tattooprobes/probes.py:

NEEDLE_MM = 0.30   # approx. single-needle/3RL line (practitioner guidance; assumption, see paper)
GAP_MM = 0.50      # gaps narrower than this are assumed to close after healing (assumption)
WORK_PX_PER_MM = 10.0  # analysis resolution after physical rescaling

def to_gray(img):
    a = np.asarray(img.convert('RGB') if hasattr(img, 'convert') else img)
    return cv2.cvtColor(a, cv2.COLOR_RGB2GRAY) if a.ndim == 3 else a

def physical_resample(gray, size_cm):
    """Rescale so that the longest side spans size_cm at WORK_PX_PER_MM."""
    h, w = gray.shape
    target = int(round(size_cm * 10 * WORK_PX_PER_MM))
    s = target / max(h, w)
    interp = cv2.INTER_AREA if s < 1 else cv2.INTER_CUBIC
    return cv2.resize(gray, (max(1, round(w * s)), max(1, round(h * s))), interpolation=interp)
Enter fullscreen mode Exit fullscreen mode

size_cm * 10 converts centimetres to millimetres, then WORK_PX_PER_MM scales the long side. INTER_AREA handles downscales and INTER_CUBIC handles upscales. to_gray accepts a PIL image or an array.

ink_mask is Otsu, or 128 when the gray standard deviation is at most 1. Ink is gray < t. _skel_widths samples a distance transform on the skeleton, doubles the radius, and divides by 10 px/mm. line_w_p10_mm is the 10th percentile of that sample. frac_sub_needle is the share under 0.30 mm. frac_tight_gap is the same width share on background inside the ink box, counted where the width is under 0.50 mm. midtone_frac is the share of gray values strictly between 60 and 200. edge_density_per_mm2 is the Canny count at 50 and 150, divided by WORK_PX_PER_MM and by area in mm².

PROBE_SIGNS is fixed a priori. line_w_p10_mm is +1. frac_sub_needle, open_ends_per_cm, frac_tight_gap, midtone_frac, border_std, edge_density_per_mm2, specks_per_cm2, and edge_transition_mm are −1. The composite is the unweighted mean of signed robust z-scores, median and IQR over the pool. edge_transition_mm is post hoc, from after the first degradation run, and probes_preregistered omits it.

from tattooprobes.probes import compute_probes
from tattooprobes.simulate import retention

features = compute_probes(img, size_cm=5.0)
kept = retention(img, size_cm=5.0, spread_mm=0.30, seed=0, noise=0.0)
Enter fullscreen mode Exit fullscreen mode

features is a float dict. kept holds skel_f1 and gap_survival. We call both at 3, 5, 10, and 15 cm. When spread_mm > 0, simulate_on_skin runs cv2.GaussianBlur(g, (0, 0), spread_mm * WORK_PX_PER_MM) and adds noise only if noise > 0. The runs below pass noise=0.0.

Excerpt from tattooprobes/simulate.py:

    deg = simulate_on_skin(img, size_cm, spread_mm, seed, noise)
    # v0.3 fix: threshold the simulated image with the SAME Otsu threshold as the reference, and
    # compare skeletons with a tolerance expressed in mm (0.2 mm) instead of a fixed 5-px window,
    # which previously made the match stricter at larger physical sizes (resize artifact).
    from skimage.filters import threshold_otsu
    t0 = threshold_otsu(ref) if ref.std() > 1 else 128
    m1 = deg < t0
    s0, s1 = skeletonize(m0), skeletonize(m1)
    k = 2 * int(round(TOL_MM * WORK_PX_PER_MM)) + 1
    tol = cv2.dilate(s1.astype(np.uint8), np.ones((k, k), np.uint8)) > 0
Enter fullscreen mode Exit fullscreen mode

TOL_MM is 0.2. F1 is the harmonic mean of the two dilated skeleton overlaps. Gap survival is the share of reference background inside the ink box that is still background after the blur.

Probe pipeline (Illustration)

How the rows got onto disk

HPDv2 train, prompts matching "tattoo": 2,010 pairs, 327 prompts, 1,340 images, stream-extracted from a 31.7 GB tar in 23 minutes, about 60 MB kept on disk. ImageRewardDB train+val tattoo prompts: 370 rows. 47 images are 0-byte in the upstream zips, leaving 323 usable images, 809 within-prompt rank pairs, and 51 prompts. Pick-a-Pic v2 came from a liuhuohuo2 parquet mirror of 672 shards, metadata only: 2,608 tattoo rows, 1,687 labeled and different pairs, 185 prompts. Images for that pass were not fetched. The license is unverified. The original card is gone, and the mirrors say MIT. Drozdik/tattoo_v0 is CC-BY-NC-SA-4.0, research only: 120 designs sampled, 6 defect types at 2 levels, 1,560 images.

Labels are a keyword rule plus CLIP ViT-B/32 zero-shot (open_clip ViT-B-32, pretrained openai). HPDv2: design 334, on_skin 87, incidental 674, other 146, ambiguous 99. ImageRewardDB: design 48, on_skin 12, incidental 238, other 18, ambiguous 7. tattoo_subject is design plus on_skin, which is 463 pairs and 119 prompts on HPDv2, and 148 pairs and 9 prompts on ImageRewardDB. A spot-check of 24 random design images looked tattoo-like for ~21/24 on HPDv2 and 24/24 on ImageRewardDB.

Per-dimension rows below are the alignment after the reward re-score: 528 pairs and 125 prompts. The composite stays on 463 pairs and 119 prompts.

Agreement, holdout, and the spread ladder

Pairwise accuracy is the rate at which the scorer orders a pair as the human label did. On the HPDv2 tattoo-subject per-dimension table, n_pairs is 528 and n_prompts is 125. Holm p is 0.0065 on every row.

Scorer Accuracy A A−0.5 Cohen h Holm p sig
frac_tight_gap@5 0.335 −0.165 −0.34 0.0065 True
edge_density_per_mm2@5 0.335 −0.165 −0.34 0.0065 True
midtone_frac@5 0.369 −0.131 −0.26 0.0065 True
line_w_p10_mm@5 0.428 −0.072 −0.14 0.0065 True
CLIPScore 0.682 +0.182 +0.37 0.0065 True
PickScore 0.723 +0.223 +0.46 0.0065 True
ImageReward 0.684 +0.184 +0.38 0.0065 True
HPSv2.1 0.790 +0.290 +0.62 0.0065 True

probe_composite@5 is the other alignment: 0.305 [0.256, 0.353] on 463 pairs and 119 prompts.

Per-dimension agreement with human preference, HPDv2 tattoo-subject

Fusion fits LogisticRegression on winner-minus-loser features after a random side flip, with a StandardScaler and GroupKFold on prompt id. HPDv2 rows in results/fusion_lopo.csv are GroupKFold(prompt,k=5). A literal leave-one-prompt-out loop will not match that file. Preregistered probes, with edge_transition_mm excluded, score 0.708 [0.667, 0.750]. Rewards score 0.790 [0.748, 0.827]. Rewards plus probes score 0.767 [0.729, 0.803]. Probes alone beat chance in that fit only with inverted coefficient signs.

Fusion LOPO accuracy

In the default simulator, noise stays at 0. At σ = 0, mean skeleton F1 is 1.000 at every tested size. At σ = 0.30 mm, mean skeleton F1 by print size is 3 cm 0.477, 5 cm 0.563, 10 cm 0.668, 15 cm 0.733. At 5 cm and σ = 0.30 mm, Spearman ρ is −0.824 for frac_tight_gap vs gap survival, −0.599 for edge density vs gap survival, and +0.507 for line_w_p10_mm vs skeleton F1.

Minimum print size vs ink spread

What we would change in the next eval

Fusion at 0.708 [0.667, 0.750] tracks the human bit, and the probes clear chance in that fit only with inverted coefficient signs. The signed composite on 463 pairs is 0.305 [0.256, 0.353]. Log PROBE_SIGNS for the millimetre direction. Fit the human bit for this preference label. A training loss built from the signed mean follows the 0.305 direction. line_w_p10_mm is the only preregistered key with sign +1, so a fatter Otsu mask, whether from thickening or from blur, moves that coordinate the way the heuristic favors.

HPSv2.1 was trained on HPDv2, and ImageReward was trained on ImageRewardDB, so those agreements are in-distribution leakage.

ImageRewardDB tattoo-subject has 9 prompts. The Pick-a-Pic design slice has 11 prompts. Probes are not Holm-significant on either slice, so both stay out of the 528-pair table.

frac_tight_gap and gap survival both read background inside the ink box: widths already under 0.50 mm, and pixels still background at σ = 0.30 mm. ρ = −0.824 is partly circular by construction. Edge density vs gap survival is ρ = −0.599. line_w_p10_mm vs skeleton F1 is ρ = +0.507.

The 5 px window in the excerpt made the match stricter at larger physical sizes. TOL_MM = 0.2 keeps the band in millimetres. At σ = 0 the published mean skeleton F1 is 1.000 at every tested size.

The ink-spread model is an isotropic Gaussian, not validated on healed skin. tatany.app was not ranked.

Reproduce

git clone https://github.com/Stark-Will/tattoo-probes
Enter fullscreen mode Exit fullscreen mode

The archive is doi:10.5281/zenodo.23292900. Install requirements.lock.txt, keep the CPU torch wheel, and leave the seeds at 0. Entry points: compute_probes(img, size_cm=5.0) and retention(img, size_cm, spread_mm, seed=0, noise=0.0). Read per-dimension rows from the 528-pair, 125-prompt alignment, and the composite from the 463-pair, 119-prompt alignment.

Disclosure: I work on tatany.app, an AI tattoo generator; the study does not rank it or any other product.

Top comments (2)

Collapse
 
aifliproom profile image
AI Flip Room •

Really honest write-up, especially the part where the probes only beat chance with flipped signs. We hit something similar checking AI room photos: hand-made pixel metrics were great at catching one specific failure (a line that vanished, a gap that closed) but useless for ranking which image people liked more. We ended up using cheap checks as hard gates and leaving the preference call to a model. Curious if you'd try the probes as filters rather than in the score?

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

Official Platform Update

Security protocols have been updated for all developer accounts.

  • tr.ee/dev-to