Methodology
A score on this site is built from four numbers: capabilities (40%), quality of output (30%), ease of use (15%), and price & limits (15%). Here's what actually goes into each one, and just as importantly, what doesn't.
Capabilities — Can It Actually Do the Job?
This is the biggest slice of the score because it's the most basic question a review can answer, and the one marketing pages are least reliable about. We test the specific claims a tool makes — not a generic benchmark, but the actual task a reader would try. If a coding assistant claims it "understands your whole codebase," we test it on a real multi-file project, not a fifty-line demo. A feature that only works in the promo video doesn't earn points here.
Quality of Output — Good Once, or Good Every Time?
A single great result proves nothing. We run the same category of task repeatedly and look at the floor, not just the ceiling — how bad the worst output gets, not just how good the best one is. This matters more in some categories than others: a writing tool that's brilliant nine times out of ten and flatly wrong the tenth time is a bigger problem than a slightly-below-average tool that's at least consistent.
Ease of Use — The Fifteen-Minute Test
We time how long it takes a first-time user to go from signup to a result they'd actually use. A tool that requires reading documentation before it does anything useful loses points, even if the underlying model is excellent, because most people never make it past that wall. This is weighted lower than capabilities and quality on purpose — we'd rather point you to a slightly clunkier tool that's genuinely better than a smooth one that's mediocre underneath.
Price & Limits — What You Get, Not What's Advertised
The number on the pricing page is the least interesting part of this criterion. What matters is what you actually get for it: how many generations before a paywall interrupts you, whether the free tier is a real product or a seven-day countdown, and whether "unlimited" comes with an asterisk buried in the terms. We've scored technically-cheaper plans lower than pricier ones when the cheap plan's limits made it unusable for normal work.
What We Don't Score
We don't factor in a tool's funding round, its logo, or how recently it launched. A well-funded tool with a bad free tier doesn't get a pass for "being new," and a scrappy tool from a two-person team doesn't get penalized for it either. We also don't score marketing claims we haven't personally verified — if we haven't tested it, it's not in the review.
When Scores Change
Scores update whenever a vendor ships something that would plausibly move the needle: a new underlying model, a pricing change, a feature removed or added. We don't run on a quarterly schedule, because AI tools don't ship on one either. Check the "updated" date on any review before assuming a score reflects the tool as it exists today.
The Honest Limits of This System
No fixed methodology captures everything. A tool can score well here and still not fit your specific workflow, your specific writing voice, or your specific codebase's conventions. Treat the number as a strong starting point, and the review text — which explains where and why a tool earned that number — as the part that actually helps you decide.
An Example Walkthrough
Take two hypothetical writing tools scoring 8.9 and 8.6. The higher-scoring one might win almost entirely on quality of output — noticeably better structure and fewer factual slips across repeated tests — while losing slightly on price because its free tier caps out fast. The lower-scoring one might be nearly as good on quality but meaningfully faster to start using with zero setup. Neither number tells you which one fits a specific afternoon of work better than the review text does; that's why we write the review at all instead of just publishing a spreadsheet.