Market assessment · Advise

Agentic AI pricing: where the inference efficiency gain stops

Nvidia's own generation claims compound to roughly 450x more throughput per megawatt since Hopper. Over the same window the flagship token price sheet moved 3x, and its top end went up. The gain is retained where the meter sits, and agent workloads consume enough extra tokens to absorb what does get through.

26 PAGES EVIDENCE-LABELED FREE TO READ
Commission an assessment

The efficiency gain is real. Almost none of it reaches the buyer.

This report follows one claim from the silicon to the invoice. The claim is that a new generation of inference hardware makes agentic workloads dramatically cheaper to serve. Nvidia states that GB300 NVL72 delivers up to 15x better throughput per megawatt than Hopper, and that Vera Rubin NVL72 delivers up to 30x better than GB300. Multiply those and you get about 450x since Hopper, a figure Nvidia does not itself publish. Over a comparable window Anthropic's flagship went from 15 and 75 US dollars per million input and output tokens on Opus 4.1 to 5 and 25 on Opus 5. That is a clean 3x. The top of OpenAI's published sheet moved the other way, with GPT-5.5 Pro listing at 30 and 180.

Three things about this category's published numbers need discounting before any of them are used. Every headline efficiency figure is a vendor figure, measured by the vendor, on a configuration the vendor selected, and in two of the four cases the vendor does not disclose the batch size or concurrency behind it. The multiples are quoted against different baselines on different benchmarks, so they cannot be compared without saying so out loud. And the unit of account itself is unstable: Anthropic's own pricing documentation states that Claude 4.7 and later models use a newer tokenizer producing roughly 30 percent more tokens for the same text. A 30 percent per-token price cut across that boundary is exactly cancelled at the invoice.

The gain stops where the meter sits. Microsoft's published Copilot Studio rate card bills premium reasoning tokens at 10 Copilot Credits per 1,000 tokens, and credits list at one cent each, which converts to 100 US dollars per million tokens, charged on top of a separate per-event feature rate for the same operation. Frontier model list output pricing at the same moment runs 20 to 25 dollars. That gap, between what the model costs and what the orchestration layer bills for calling it, is where a hardware improvement of two orders of magnitude is being retained. Gartner reaches a similar place from a different direction, arguing that agent consumption grows faster than token prices fall. Both are true, and the pricing argument works even holding consumption flat.

Eight findings on agentic inference pricing and where the gain is retained

Four hardware positions are in scope: the SemiAnalysis AgentX benchmark and its methodology, Nvidia's Vera Rubin NVL72 throughput-per-megawatt claims, the Nvidia Groq 3 LPX token-generation claims, and Intel's tokens-per-watt framing for Crescent Island. Against those sit the commercial structures buyers sign: published per-million-token API pricing from Anthropic and OpenAI, published per-event agent pricing from Microsoft, Salesforce and Intercom, and the contract risk categories named by Info-Tech Research Group. Training economics, model quality, chip supply and equity valuation are out of scope.

01

AgentX measures the thing that costs money, and no two vendor claims were run on it the same way

AgentX is the first open, multi-turn agentic coding benchmark at long context, released under Apache 2.0 and built from 393 anonymised Claude Code traces with prefix reuse preserved. Its input distribution reaches a 99th percentile of 675,000 tokens against an 8.6k output p99. Fixed-sequence benchmarks in the 8k/1k class cannot see prefix reuse, KV cache lifecycle or sub-agent bursts, which is where agentic cost actually lives. Nvidia measured Vera Rubin NVL72 on AgentX and says the result is pending SemiAnalysis review. It measured Groq 3 LPX on a different benchmark, a different model and a tenth of the context.

02

Nvidia's own multiples compound to roughly 450x. The flagship price sheet moved 3x

Nvidia states GB300 NVL72 delivers up to 15x better throughput per megawatt than Hopper, and Vera Rubin NVL72 up to 30x better than GB300. Multiplied, that is about 450x since Hopper, a figure Nvidia does not publish and whose two inputs are not confirmed to share a workload. Over a comparable window Anthropic's flagship went from 15 and 75 dollars per million input and output tokens to 5 and 25. The top of OpenAI's sheet rose instead: GPT-5.5 Pro at 30 and 180, o1-pro at 150 and 600. The floor fell far faster than the ceiling.

03

The token is not a fixed unit, and the vendor controls its size

Anthropic's pricing documentation states that Claude 4.7 and later models use a newer tokenizer producing approximately 30 percent more tokens for the same text. A buyer who negotiated a 30 percent per-token reduction and then upgraded across that boundary paid exactly what they paid before, for exactly the same work. No audit clause written against a token count survives a change in what a token is unless it fixes the tokenizer version or meters in characters or requests. This is the most consequential disclosure in the evidence base and it appears as a footnote on a pricing page.

04

Copilot Studio's reasoning meter prices out near 100 dollars per million tokens

Microsoft's published rate card bills premium text and generative AI tools at 10 Copilot Credits per 1,000 tokens, and pay-as-you-go credits list at one cent each. That is 100 US dollars per million tokens, sitting on top of a separate per-event feature rate charged for the same operation, a rule Microsoft's own documentation states plainly. Frontier model list output pricing at the same moment runs 20 to 25 dollars per million. The meter bundles orchestration, connectors and governance, and no reasonable value for those closes a gap of that size.

05

One prompt now generates a dozen distinct billable events

Declaring a browser toolset adds about 6,610 input tokens to every request, not once per session, so a fifty-turn session carries roughly 330,000 tokens of tool definitions before a word of user content. Anthropic bills managed agent sessions at 0.08 dollars per session-hour on top of tokens, web search at 10 dollars per 1,000 searches, container execution at 0.05 dollars per container-hour, and cache writes at 1.25x or 2x base input. Microsoft bills 1 credit for a classic answer, 2 for a generative answer, 5 for an agent action and 10 for tenant graph grounding. Salesforce bills 0.10 dollars per action. Intercom bills 0.99 dollars per resolved outcome.

06

Every vendor in the sample charges consumption in addition to seats, not instead of them

The per-seat model has not been replaced by consumption. It has been supplemented by it. A renewal that reads as flat on the seat line can carry an unbounded second line beneath it, and the seat count no longer bounds the exposure, because an autonomous agent consumes with no human present. Gartner estimates agents require 5 to 30 times more tokens than an equivalent chatbot task, and that a task costing a chatbot one cent can cost an agent up to 1.50 dollars. Headcount-based forecasting has stopped working for this category.

07

Info-Tech names limited audit rights as a category, and no major vendor publishes a reconcilable meter

Info-Tech Research Group names five risk categories in agentic AI contracts: unclear billing definitions, hidden cost drivers, weak financial controls, vendor-controlled pricing changes, and lock-in and accountability gaps. The framing has one gap. A buyer with a cap and an audit clause still cannot verify a token count they have no independent way to produce. Caps are enforceable because the vendor meters them. Audits are not, unless the clause names the artefact being audited against. An audit right over a number only one party can produce is a right to be shown the vendor's arithmetic, not to check it.

08

The only quantified error rate in the public record is about five percent of reviewed spend, and it comes from a firm selling the audit

An audit vendor reports reviewing 34 million dollars of AI spend across 60 companies since March 2026 and identifying close to 1.7 million in mistaken overcharges, roughly 80 percent of it credited back. Discount the percentage: the firm sells the audit, the sample self-selects toward companies who suspected a problem, and no methodology is published. What survives the discounting is the taxonomy, which is checkable and matches the architecture: failed requests still charged, duplicate charges from agent retry storms, billing during provider outages, premium rates applied to older models, and orchestration errors that fan one request out to several models at once.

What the hardware claimed against what the price sheet did

These are different kinds of number, measured by different parties for different purposes, and setting them beside each other is a reasoned judgment rather than a calculation. Every figure below is a list price or a vendor statement, undiscounted, read from each company's own documentation in August 2026. Large buyers negotiate, and if enterprise discounting has moved much faster than list pricing then the size of the gap is overstated. The direction is not in dispute.

Nvidia, claimed

Up to 450x throughput per megawatt since Hopper

Two Nvidia figures multiplied: 15x from Hopper to GB300 NVL72, then 30x from GB300 to Vera Rubin NVL72. Nvidia publishes both and neither the compounded number nor a confirmation that they share a workload. Stated as up to, measured by the vendor, never independently reproduced.

Vendor claim, pending review

Anthropic, published

3x on the flagship tier

Claude Opus 4.1 listed at 15 and 75 US dollars per million input and output tokens. Claude Opus 5 lists at 5 and 25. A factor of three on both sides, over roughly the window in which Nvidia claims 450x, on the tier an agentic workload actually runs on.

15 and 75 down to 5 and 25 USD

OpenAI, published

The floor fell and the ceiling rose

GPT-5.6 Luna lists at 0.20 input and 1.20 output per million tokens, so the cheap tier really did collapse. GPT-5.5 Pro lists at 30 and 180, and o1-pro at 150 and 600, above the March 2023 GPT-4 output price. Cheap-tier deflation is close to irrelevant to an agent holding 675,000 tokens of context.

0.20 to 180 USD across one page

Microsoft, published

Roughly 100 USD per million tokens on the premium meter

Copilot Studio bills basic text at 0.1 credits per 1,000 tokens, standard at 1.5 and premium reasoning at 10, with credits at one cent each. That is 1, 15 and 100 dollars per million, and the premium tier is charged on top of the per-event feature rate for the same operation.

Four to five times frontier list

Throughput per megawatt and cost per token are different quantities, and the gap between them is capital. Analyst estimates put a GB200 NVL72 rack near 3 million US dollars, a GB300 NVL72 rack between 3.7 and 6.5 million depending on whose estimate, and a VR200 NVL72 rack at roughly 7.8 million on a Morgan Stanley bill of materials. The near-twofold spread between two estimates of the same rack is itself the finding: nobody outside the supply chain knows this number, so no buyer can independently check a cost-per-token claim built on it. A 30x throughput-per-megawatt part running at 20 percent utilisation is not a 30x cost improvement, and none of the published claims state a utilisation assumption. The efficiency gain is real for anyone whose binding constraint is power, and much smaller for anyone whose binding constraint is capital or utilisation, which is most operators outside the largest clouds.

Four clauses, in priority order

Define the billable event, including the tokenizer version and a notice requirement before it changes, because a tokenizer producing 30 percent more tokens for the same text will silently reverse a 30 percent price concession. Enumerate every meter in a schedule and not just the token rate, because the ones you did not name are the ones you cannot dispute: cache writes, tool definition overhead, server tool calls, session-hours, container-hours, data residency and speed multipliers. Take the hard cap. Then negotiate the audit clause against a named artefact rather than against a token count, since the count is a number only the vendor can produce.

The hard cap is not a hypothetical ask. Microsoft's published enforcement policy triggers at 125 percent of prepaid capacity, disables custom agents, lets in-flight conversations complete, notifies administrators by email and in the admin centre, and allows per-agent monthly limits to be set below the contractual ceiling. It fails closed, it fails gracefully, and it names a human. One large vendor already ships it, which closes the usual objection that caps are operationally hard. A vendor declining a hard cap in 2026 is declining a control a competitor ships by default.

Two negotiating levers are undervalued. Price holds matter more than discounts here, because the rate card has more lines than the contract and any one of them can move, so a hold on all meters for the term is worth more than a few points off the token rate. And the right to exit at the old price if any meter changes converts a unilateral pricing right into a bilateral one.

Every claim carries its evidence

This isn't a vendor summary. Every sentence is labeled by what stands behind it: verified fact, vendor claim, third-party estimate, my assessment, hypothesis, or scenario. Sources are numbered and clickable. Forward-looking sections use scenarios with observable tripwires, not forecasts. It's the same method behind every market assessment I write.

Nvidia Says 450x. Your Invoice Says 3x.

Twenty-six pages, built from public sources with no client brief and no interviews. Read it in the browser or take the PDF.

This is real, published work. Commission one for your decision.

Each report here answers a real question, directed and researched against public sources and evaluated against a stated assumption, then delivered as Word and PDF. If you're weighing a platform, sizing a category, or defending a number to a board, tell me the decision behind it and I'll tell you honestly whether a report is the right tool.

Commission an assessment

See more reports →