Testing methodology
We test tools against a finished production job, not a feature checklist. The unit might be one accepted video minute, one completed episode, one usable image, or one stable regional test profile. That denominator stays visible so a cheap input cannot hide expensive retries.
Evidence hierarchy
| Label | What it means | How we use it |
|---|---|---|
| Measured | Observed directly in our own run, logs, hardware telemetry, or invoice. | May support a firm conclusion. |
| Calculated | Derived from measured inputs with the arithmetic stated. | Used for workload projections and break-even comparisons. |
| Estimated | Based partly on an assumption, vendor rate, or non-identical workload. | Shown as a range and never presented as a benchmark. |
What a useful test records
We record the hardware or service tier, model and quantization where relevant, input dimensions and duration, output dimensions, run time, billable usage, retry count, and test date. We also preserve failure notes because unusable outputs and manual recovery often dominate the real cost.
Comparison rules
We normalize options to the same finished workload whenever possible. If two paths differ in setup time, control, quality, or turnaround, those differences stay in the conclusion rather than being flattened into one misleading dollar figure. Vendor list prices can provide context, but they are not labeled as our measured results.
Updates and limitations
AI products and prices change quickly. Each article carries a tested or updated date. A single production setup cannot represent every geography, prompt, model, or skill level; readers should reproduce the smallest relevant test before committing a large workload. Material errors can be reported through the contact page.