DEV Community

Alex
Alex

Posted on • Originally published at saas.pet

I tested 213 AI tools so you don't have to — here's what I learned after 3,943 days

I started tracking AI tools in mid-2015. Today it means 213 reviews on saas.pet — every one paid for out of my own pocket, used in real work, rated on what actually shipped.

The thing nobody tells you about the AI tools boom is that most of them are noise. Of the 213 tools I've tested on saas.pet, only about 30 have stayed in my daily workflow. The rest either got replaced by better alternatives, turned out to be vaporware, or solved a problem that didn't exist.

This is what I learned after 3,943 days of testing.

The first lesson: usage beats features.

When I started saas.pet, I rated tools on capabilities — features, model size, benchmark scores. After 200+ reviews, I learned the only metric that matters is whether you actually open the tool tomorrow morning. The fanciest AI agent in the world is useless if you switch back to ChatGPT for real work.

This is why my reviews have a days_used field. A tool with 4/5 rating after 90 days of daily use is more trustworthy than a 5/5 rating after a 7-day trial. The longer I lived with the tool, the more I learned about its actual weaknesses.

The AI tools I kept coming back to all share one trait: they got out of my way. They don't require a prompt engineering course. They don't crash on long contexts. They integrate with the way I already work. The tools that fail are the ones that force you into their workflow.

The second lesson: data labeling platforms matter more than you think.

When you use ChatGPT, you're seeing the output of a labeling pipeline that took months to build. Most people never think about this. After testing Labelbox and Surge AI — both data labeling platforms that power frontier AI labs — I realized the quality of your model's training data matters more than the architecture. Labelbox claims 90% of leading US AI labs are partners. Surge AI counts Anthropic, OpenAI, and Google DeepMind among its customers. These are the companies quietly deciding what your AI is good at.

I haven't personally run labeling projects on either — they're enterprise infrastructure with custom pricing — but reading their public case studies was eye-opening. If you're building production AI, your data labeling pipeline is your moat, not your model.

The third lesson: GPU clouds are not interchangeable.

When I needed to fine-tune a model last year, I assumed all GPU clouds were the same. CoreWeave changed that assumption. CoreWeave is a GPU-specialized cloud provider that Gartner named a Visionary in 2026 Cloud AI Infrastructure — purpose-built for AI workloads rather than retrofitted from general-purpose compute. If you've ever waited 20 minutes for a GPU to spin up on a generalist cloud, you understand why specialist providers matter.

I haven't personally run production workloads on CoreWeave either (it's enterprise-scale pricing), but the public engineering blog posts reveal a different design philosophy: shared GPU scheduling, sub-second cold starts, and pricing aimed at GPU-heavy AI rather than web hosting. The generalist clouds are catching up, but the specialist clouds are still ahead for serious AI work.

The fourth lesson: the tools you don't pay for are the ones you forget.

I have a policy on saas.pet: I pay for every tool I review, out of my own pocket. No free accounts from vendors, no "reviewer program" perks, no sponsored reviews. This costs me about $50/month on top of the tools I use anyway (you can see the full breakdown in my AI tools cost analysis).

Why does this matter? Because free accounts come with usage limits that distort testing. When a tool gives you 100 free generations and you're on generation 95, you're not testing the tool — you're rationing. When you pay, you test the way you'd actually use it. This is why my reviews focus on the experience of someone who depends on the tool daily, not someone optimizing for free tier credits.

The eight tools I keep coming back to.

After 213 reviews, here are the 8 I keep open in browser tabs every single day.

ChatGPT is still my default for quick reasoning tasks. I have ChatGPT Plus ($20/month) and use it probably 5-10 times per day. The context window improvements and image generation in GPT-4o make it the Swiss Army knife of AI tools — not the best at any one thing, but consistently good at most things.

Claude is what I switch to when I need longer thinking. Claude Sonnet 4.5 reads my whole codebase before answering — no other tool does this as reliably. For code review, document analysis, or anything where I need the model to actually read what I gave it, Claude wins. I have Claude Pro ($20/month) and use it probably 10 times per day for serious work.

Cursor replaced VS Code for me 8 months ago. The inline AI editing, the multi-file context, the way it learns my coding style — nothing else comes close for daily development. I pay $20/month and consider it the best $20 I spend on tools. Cursor is the first AI coding tool that felt like a real IDE rather than a chatbot bolted onto a text editor.

Midjourney is still my image generation default. I have Midjourney Standard ($30/month) and have generated probably 10,000+ images for saas.pet, blog posts, and product mockups. Midjourney v7's text rendering and prompt adherence are the best in the industry. The Discord-first workflow still feels weird to new users, but the quality justifies it.

Perplexity is what I use when I need to verify a fact or research a topic. Pro Search pulls from multiple sources and shows me the citations inline — no other AI search tool does this as transparently. I have Perplexity Pro ($20/month) and use it instead of Google for most research queries now. The Comets browser integration makes it even better for sustained research sessions.

Hallmark is the surprise on this list. Hallmark is a 5K-star design skill for Claude Code and Cursor that enforces UI consistency across generations. If you've ever used AI to generate UI and gotten 5 different button styles in a row, you understand why Hallmark matters. It's free, open source, and solved a problem I didn't realize was solvable.

Granica AI is in a category I didn't know existed until I needed it. Granica is a data preprocessing platform that compresses and deduplicates training data — the kind of infrastructure you only think about when your model is choking on its own dataset. I haven't run training pipelines on Granica (it's enterprise infrastructure), but I learned from the documentation that data preprocessing is where most AI projects either scale or stall.

The fifth lesson: don't trust a single benchmark.

I see reviews that say "Tool X scores 19/20 on coding benchmarks" and I know immediately the reviewer ran the benchmark once and called it a day. After testing this many tools, I've learned that benchmarks measure what the benchmark designer thought to test. Real work surfaces the gaps.

The pattern in my reviews: I score a tool 4/5 if I keep using it daily after a month. I score 3/5 if I use it weekly. I score 2/5 if I drop it after a few weeks. This is more honest than running benchmarks, because it captures what actually survives in your daily workflow.

How I keep this sustainable.

Running saas.pet costs me about 3 hours per day for testing and writing. The reviews themselves follow a consistent structure — what it does, what I use it for, what works, what doesn't, pricing, who should buy. This pattern is documented at my best AI tools pillar guide, which gets updated quarterly.

If you want to browse the full collection of 213 reviews, saas.pet has them all — searchable by category, rating, and days used. The find page lets you filter by use case. The best-of guides compile the top picks by category.

The honest takeaway.

The AI tools boom is real, but most tools are forgettable. The few that survive in daily use share one trait: they solved a specific problem better than the alternative, then got out of my way. I keep paying for those tools because they earn it every day.

If you're picking AI tools for your own work, my advice is simple: pay for the tool yourself, use it for a month, and see if you still reach for it after the novelty wears off. The tools that survive that test are the ones worth recommending.

I keep a running list at saas.pet/today.html — every tool I've reviewed in chronological order, with my honest rating after living with it for the long haul.

Originally published on saas.pet

Top comments (0)