DEV Community

Cover image for How I Got Benable's List Optimized Badge (Using Claude to Hit All 8 Criteria)
Phil Rentier Digital
Phil Rentier Digital

Posted on • Originally published at rentierdigital.xyz

How I Got Benable's List Optimized Badge (Using Claude to Hit All 8 Criteria)

Benable, the app that pays you to share stuff you use, emails me: my list just earned the List Optimized badge πŸ†. Straight to the point, no warm-up, with a link to their 8 algorithm criteria. Decent of them to share. Except the public criteria say things like "high quality descriptions" or "optimized to be discovered," and that tells you exactly nothing about what to actually write in an item note to move the score.

So the real question is what actually flips a list from unoptimized to optimized, once you strip the marketing gloss off the email. The badge isn't a switch you flip by checking a box, it's a composite score, recalculated, and Benable never tells you the weights.

I took the 8 vague criteria and turned them into testable Claude prompts, one by one. Verification instead of guesswork.

The Email That Started It

The email landed on a Tuesday, tied to one specific list: "Etsy Seller Tools I Actually Use for SEO, Pinterest & More Sales." Benable's team congratulates you, drops a badge icon, and links to their tips page. On that page: 8 bullet points. Number of items. Quality of descriptions. Title structure. Image density. Section descriptions. Link type mix. A couple more phrased just as loosely.

Nowhere does it say how to measure any of it before you publish. No threshold, no example of a passing description versus a failing one, no word count target. You're supposed to intuit "high quality" from a sentence fragment written for a general audience of list creators, most of whom are not going to sit down and reverse-engineer a scoring function.

That gap is where this gets interesting. A vague criterion you can't test is a criterion you can't fix on purpose, you can only stumble into it. And stumbling is exactly what I'd been doing until the badge showed up.

The Badge Is a Score, Not a Switch

Best I can tell from watching the number move, the badge tracks a composite score, and on this list it moved between 3.78 and 4.82 over a few sessions. Not a boolean you cross once and forget, a number that goes up when you improve specific dimensions and (I'd bet, though I haven't tested it directly) probably drifts back down if you neglect the list long enough.

It's a bit like a hidden stat bar in an RPG. You can watch the number tick up after you do something, but the game never shows you the formula, just the result.

Benable's 8 public criteria are the inputs. What's missing is the weighting. Does a stronger title move the needle more than better item descriptions? Does image density matter as much as link type diversity? No idea, and Benable isn't telling. Maybe I'm reading too much into the silence, could be it's genuinely just a rough heuristic and not a precision-tuned model. Either way, you're optimizing against a black box with a public list of ingredients and no recipe.

That's the setup that made me stop guessing. If the keyword difficulty gap behind Benable payouts taught me anything about this platform, it's that the stuff Benable tells you outright is usually real, just incomplete. Same pattern here: 8 real signals, zero calibration data.

A checklist you can't test isn't a checklist, it's a suggestion.

Turning 8 Vague Criteria Into 8 Testable Prompts

This is the part that actually moved the score. For each vague criterion, I wrote a Claude prompt that either audits the current state or rewrites it against a concrete standard. 3 examples, the ones that did the most work.

Item descriptions too short or too salesy:

Here are 15 item descriptions from my Benable list. Flag every one under
200 characters or over 400. For the ones that read like ad copy
("game-changing," "must-have," "you won't believe"), rewrite them in
200-400 characters using specific, lived detail: what I actually used it
for, what broke, what surprised me. No superlatives unless I gave you a
number to back it up.
Enter fullscreen mode Exit fullscreen mode

Generic title, no curator voice:

My list title is "Best Etsy SEO Tools." Rewrite it in first person,
implying I actually use these tools daily, not that I researched a
category. Keep it under 60 characters. Give me 5 variants ranging from
blunt to slightly opinionated.
Enter fullscreen mode Exit fullscreen mode

Empty or missing section descriptions:

Here's my list structure with section headers and item counts per
section. Some sections have no description at all. Write a 1-2 sentence
description for each empty one, explaining what a reader gets from that
specific section, not a restatement of the header.
Enter fullscreen mode Exit fullscreen mode

Some of these landed on the first try. The title rewrite worked in 1 pass, obviously, the score bump was visible within a day. The description rewrites took 3 rounds because Claude doesn't have eyes on Benable's actual scoring function, just proxies: length, keyword presence, structural markers. It's optimizing what it can see, and what it can see isn't the same thing Benable is measuring.

If you want the broader version of this same idea, checking things before you ship instead of after Benable's badge email tells you what's wrong, I packaged the general framework as the Private API Checklist.

Somewhere in the middle of round 2, my Convex dashboard lost track of a webhook for reasons I still can't explain, and I burned 40 minutes debugging that instead. Unrelated. Just how the week went. The kind of dΓ©tour that makes you question if you should have just manually updated the descriptions instead of building a whole prompt system, but here we are.

That mismatch, prompt against proxy instead of prompt against target, is the whole game here. You're never optimizing the real thing directly. You're optimizing your best guess at what correlates with it, then checking the badge score to see if the guess held, like debugging with print statements when the debugger's sitting right there, unused. It's the same discipline behind the prompt contracts framework I use to avoid guessing, define the target before you write a single prompt, then check the result against something real instead of trusting the output on vibes.

Before and After, With the Actual Numbers

TITLE "The Benable Score Climb" + subtitle "3.78 to 4.82 across three passes". Metaphor: a mountain trail with 3 rest stops, each marked with a flag showing a score. Style: engineer blueprint, thin white lines on navy background, technical annotation marks. Palette: navy #14213D, amber #FCA311, muted red #C1121F, cream #F5F1E8. Content: 3 checkpoints labeled UNOPTIMIZED (score 3.78, 8 items, 0 images, 100 percent deal links), NEAR-OPTIMIZED (score 4.32, 12 items, 2 images per item, 70 percent deal links), OPTIMIZED (score 4.82, 15 items, 3 to 4 images per item, 50 percent deal links). Highlight: the OPTIMIZED checkpoint flag glows amber with a small badge icon next to it. Legend: none needed. Footer: Β© rentierdigital.xyz. NOT flat corporate vector, NOT minimalist tech startup aesthetic.


Benable Score Optimization Journey: Three Performance Milestones

Here's the progression, and it wasn't a straight line.

The score started at 3.78, with 8 or 9 items, zero images, and close to 100 percent affiliate deal links, which in hindsight was probably dragging the link type mix criterion into the floor the whole time. The first pass of prompts got it to 4.32, items up to 12, images added on most entries, descriptions rewritten, and that felt like a finish line until I checked the number again and realized it wasn't one. The second pass, targeting the title and the section descriptions specifically, plus swapping roughly half the deal links for tools I actually use with no commission attached, is what pushed it to 4.82, and that second push took longer than the first because none of the remaining gaps were obvious from reading the list myself.

The line item that surprised me most: link type mix. I'd assumed Benable wanted more links, period. What actually moved the score was variety, deal links mixed with plain tool recommendations, not volume. That's the kind of thing a tips email skips because it's not a clean bullet point, it's a ratio you only find by testing.

What This Actually Cost

3 sessions, spread over about a week and a half, is what it took to get from 3.78 to 4.82. The title and image fixes were fast. The description rewrites and the link mix rebalancing took iteration, because Claude was working off proxies the whole time, never off the actual score itself.

What I still don't know: the exact weight Benable puts on each of the 8 criteria. Whether the badge can drop back down if a list goes stale (my working assumption is yes, based on nothing but how composite scores usually behave, so take that with a grain of salt). And the no_index status Benable applies to unoptimized lists doesn't have a workaround through the API, I checked. You optimize the list the normal way or your list stays invisible to search, full stop. No side door, no admin flag to flip, just the grind, like farming a boss with no drop table you're allowed to see.

3 sessions and a handful of prompt rewrites bought a score that only tells you it moved, never why, never by how much.

Sources

  • Benable's List Optimization Tips, the official criteria page linked in the badge email
  • Email from the Benable team, August 7, 2026, badge announcement for the Etsy Seller Tools list

This post may contain affiliate links. If you click them, I might earn a small commission (costs you nothing, and helps me keep shipping quality articles every day for your reading pleasure).

Top comments (0)