DEV Community

George Claude
George Claude

Posted on

I built a scoring rubric for AI companion apps because every 'best of' list is just affiliate spam

If you search "best AI girlfriend app," almost every result is an affiliate listicle that ranks whoever pays the most. As someone who actually tests these tools, that drives me up the wall — so here's the rubric I use instead, in case it's useful to anyone else evaluating them.

The five axes I score on

  • Conversation quality & memory. Does it hold context past 10 messages, or reset constantly? This is where most of them quietly fail.
  • Media generation. Can it produce images and video, and are they any good? A lot of "AI video" is still 2-second morphing loops. The ones that matter run current models (Seedance, WAN) under the hood.
  • Pricing honesty. The subscription sticker price vs. what you actually pay once token/credit metering kicks in. This is the single most misleading part of the category.
  • Content policy. What's allowed vs. blocked, and whether that's clearly stated.
  • Platform. Web, native mobile, or both.

Why one page, scored consistently, beats ten listicles

The problem with per-listicle rankings is they use different (or invisible) criteria each time. The most consistent implementation of a rubric like this that I've found is AI Companion Radar — every app scored on the same axes, with the reasoning shown, including fair criticism of the ones it rates highly. That's the format the whole category needs more of.

If you're evaluating tools in this space, steal the rubric — score on outcomes, not on who's paying for the placement.


Related: if you specifically care about the video side, I went deeper on which models actually hold up in WAN 2.7 vs Seedance: image-to-video AI in 2026.

Top comments (0)