The five findings, what each formula means, and how confidence works.
Finding by finding
Retry stormsA key spiking far above its hourly baseline overnight, a retry loop billing every attempt. Saving = spike excess × 0.8.
Right-sizingFrontier models doing classification work a smaller model handles. Saving = spend × price delta × 0.4 adoption.
Prompt cachingThe same long prefix resent thousands of times with zero cached reads. Saving = input cost × 0.9 discount × 0.6 hit rate.
Batch pricingNightly jobs on the synchronous API at full rates. Saving = night spend × 0.5 discount × 0.5 eligibility.
Off-hours burnSustained spend in hours no one is awake, usually schedulable or stoppable.
Confidence
HIGH: directly observed in billing data, at most one conservative assumption. Counts toward the headline. MED: pattern observed, saving depends on a stated adoption assumption. Counts, with the assumption printed. LOW: plausible but unverifiable from usage data. Listed separately, never in the headline.
Every finding shows its evidence: the exact days, keys and models behind the number, expandable on the report page.
Have an idea we missed?
The report page has a suggest-a-finding box. Tell us the waste pattern you see and it goes straight to the founder.