DEV Community

zhouhua xiao
zhouhua xiao

Posted on

How I Evaluate AI Tools: A Small, Repeatable Test

I used to evaluate AI products by their first demo. If the answer sounded fluent, the interface looked polished, and the generation felt fast, I was impressed. After a few days, however, the real problems usually appeared in the details. Could the tool follow a fixed format? Could I trace an answer back to its sources? Would it admit uncertainty? Would the free allowance cover one complete job rather than one impressive preview?

My current process has four steps. First, I define a real task with a clear boundary. Second, I prepare fixed inputs so the test is comparable. Third, I record time to completion, rework, error types, and final usability. Fourth, I keep product claims, personal impressions, and checkable evidence in separate notes. The result is not a vague statement that a tool is “amazing” or “bad.” It is a practical description of where the tool fits and where it stops fitting.

I put the findings into four boxes: usefulness, traceability, limits, and cost. This is also why the Maker Jury approach makes sense to me. Its public methodology asks reviewers to define the user, job, inputs, comparison basis, product version, account tier, and stopping conditions before testing. New public work does not force everything into a single numeric total; it records the task, observations, measurements, limitations, and disclosures.

The historical records on the homepage are especially useful as examples. The NotebookLM record emphasizes bounded source work and traceability. The Lovable record describes a successful publication of a bounded internal tool, while also noting a long first run and unclear authorization waiting. The Framer record describes a polished, editable, responsive site, alongside cautions about a recoverable connection error and confident claims. That is closer to the answer I need than a leaderboard: what can it do, and where will it make me do the work again?

Top comments (0)