Pricing

Plans and pricing

Four plans, built to compound. Each one takes you further into your spend and returns more of it — from a first look at what is there to continuous measurement that keeps it from coming back.

Tier 1

Scan

Free

The fastest way to see the scale of what your workload is carrying.

  • Whether prompt caching is currently saving or costing you money
  • Your cache hit rate measured against the break-even rate for the cache tier your traffic is actually using
  • What the same traffic would cost on the other cache tier
  • Runs from your provider's usage export — token counts only, no prompts, no outputs, no dollar figures, nothing for your engineers to instrument
  • One page back, free, no commitment
Request a scan

Tier 2 · Start here

Cost baseline

$1,500 one-time

For teams who need the number itself — precise, sourced, and defensible in a board or budget conversation.

  • Your cost per successful request, measured not estimated
  • Full breakdown of where the spend concentrates
  • Which savings categories your workload supports, sized
  • The measurement method behind every figure, so your team can re-run it
  • Delivered in days, credited in full against a later audit
Get your baseline

Tier 3

Full audit

Scoped fixed fee

The complete picture: every finding, sized where your data supports it, in the order to ship them.

  • Every finding sized where your data supports it — the rest named, with the exact variables they still need
  • Root cause and a concrete fix for each one
  • Ranked implementation sequence your engineers can ship against
  • The quality risk of each change and what to measure to catch it — a saving that cost output quality is not a saving, it is an unmeasured trade
  • Findings ruled out are reported with the evidence, not left unmentioned
  • Baseline fee credited in full toward the audit

The fee is scoped to how much traffic there is to audit. More traffic is more work to examine.

Discuss an audit

Tier 4

Ongoing assurance

Monthly scaled to spend

Re-runs the same measurements on a schedule, so a drift in your traffic shows up as a number instead of a surprise on your bill.

  • Continuous measurement against your established baseline
  • The same checks re-run on your cadence, so a regression shows up on the next run rather than at quarter close
  • Vendor price cards re-read on a fixed schedule and re-dated, so the figures you were given stay anchored to what the vendor currently charges
  • New models and keys appearing in your export are measured on the same basis as the baseline
  • Quarterly review with your engineering and finance leads
  • Priority access for implementation support
Discuss assurance
Compare

What is included at each tier

 ScanBaseline AuditAssurance
Categories of inefficiency identifiedyesyesyesyes
Cost per successful requestyesyestracked
Spend breakdown by sourceyesyestracked
Findings sized in dollars, where your data supports itpartialyesyes
Root cause and fix per findingyesyes
Ranked implementation sequenceyesyes
Quality verification planyesyes
Regression monitoring on deployyes
Vendor price change assessmentyes
Quarterly reviewyes
Detail

How pricing works

How audits are quoted

Audit scope is set by the size and complexity of what we are measuring — how many workloads, how many vendors, and how much depth the traffic supports. We quote after the baseline, when both sides can see exactly what is there, so the fee matches the return rather than a bracket on a page. Fixed in writing before work starts.

Your baseline fee comes off the audit

Move from a baseline to a full audit and the entire $1,500 is credited against the audit fee. The baseline is the first stage of the same measurement, so you are effectively starting the audit at a discount — and arriving with the groundwork already done, which is why audits that follow a baseline move fastest.

What assurance is scaled against

Monthly cost is set against the spend under measurement, so the plan scales with you as your workloads grow. Agreed at engagement and reviewed annually — never adjusted mid-term.

Implementation support

The audit is written so your engineers can ship every change themselves — the specific change, the expected effect, the verification. Where you would rather move faster, we implement alongside your team, scoped separately, with the same before-and-after measurement confirming each change landed as sized.

What we commit to

Every finding arrives with the measurement behind it, the dollar figure its published formula gives on your measured inputs, and a fix your engineers can implement and verify. The baseline is agreed in writing before we start, so the result is checkable against your own invoices rather than taken on trust.

Start

Start now, at any tier

The scan is free and takes almost nothing from your side. Send your provider's usage export, a self-serve download from your console. It is token counts only: no prompts, no outputs, no dollar figures, nothing for your engineers to instrument. An Anthropic export runs through our tool. Any other provider's export is read by an analyst. If you already know you want the number, start at the baseline instead.

Send your provider's usage export. We come back with one page: whether prompt caching is currently saving or costing you money, your cache hit rate against the break-even rate for the tiers your traffic actually uses, and what the same traffic would cost on the other tier.

Get started   or go straight to a baseline →