Enterprise AI budgets built on falling token prices will not break on price. They break on variance. Per-token costs collapsed while per-task consumption rose faster, and the resulting cost per completed task is not merely higher, it is unbounded on the right tail. You cannot underwrite a business case against a cost distribution you cannot cap.
Almost every AI budget written in the last two years rests on one premise: per-token prices are collapsing, so the cost problem solves itself. The premise is true and the conclusion does not follow. Independent measurement puts the price of reaching a fixed capability milestone falling by anywhere from nine to nine hundred times a year, while EY's practice put the cost of a single customer service interaction at four cents in 2023 and one dollar twenty in 2026. Both can be true only if tokens consumed per completed task rose faster than price per token fell. That is the arithmetic nobody runs before signing a three-year agreement.
What stops deployments is not the level of cost per task. It is the shape of the distribution. A task that reliably costs fifteen cents can be funded. A task that costs fifteen cents at the median and eleven dollars at the ninety-ninth percentile, because a retry loop ran forty iterations against a tool that was timing out, cannot be underwritten at all. Fifty-nine percent of organisations report using agentic AI and nine percent report reaching autonomous multi-step workflows. Everyone inside that fifty-point gap has already cleared data quality, because they are running agents in production. Whatever stops them at step three is not what would have stopped them at step one.
Ten findings on the cost structure of agentic AI as it lands on an enterprise P&L, running from mid-2025 through 16 August 2026. Each is labeled by evidence type. Verified filings, vendor self-reported figures, analyst estimates and my own assessments are graded separately, and the evidence that cuts against the argument is included rather than dropped.
Epoch AI, tracking price against fixed capability milestones rather than model names, found the cost of reaching a given benchmark score falling between nine and nine hundred times per year. Over roughly the same window EY put one customer service interaction at four cents in 2023 and one dollar twenty in 2026, about thirty times higher. Tokens consumed per completed task rose faster than price per token fell.
DeepSeek moved V4-Pro output from 87 cents per million tokens flat to 1.98 dollars off-peak and 3.96 at peak, effective 16 August 2026, with peak windows at 01:00 to 04:00 and 06:00 to 10:00 UTC. The increase is the headline and the tiering is the news. Time-of-day pricing is what a provider introduces when serving capacity binds at peak, not when unit costs rise.
ServiceNow's index of 4,500 senior leaders names IT infrastructure at 71 percent, integration at 47 percent and data accuracy at 45 percent. Cost is absent. That is the strongest published evidence against this report's argument and it is presented as such. Surveys record what respondents will attribute, and nobody tells a surveyor that the business case did not close.
Fifty-nine percent of organisations report using agentic AI; nine percent report reaching autonomous multi-step workflows. That population has already passed the data quality gate, because it is running agents in production. What separates the fifty-nine from the nine is step count, and step count multiplies both token consumption and failure retries.
Gartner predicts more than 40 percent of agentic AI projects will be cancelled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls, in that order. The same firm projects agentic AI inside a third of enterprise software applications by 2028, up from under one percent in 2024. Those are compatible only if a large share of what ships is narrow enough to be cheap.
Uber's engineering organisation consumed its entire 2026 budget for one AI coding tool by April. The fix was a hard ceiling of 1,500 dollars per employee per month per tool. Users of advanced AI tools then more than quadrupled while total AI cost fell. If the constraint had been the price level, cheaper tokens would have fixed it. A ceiling fixed it, which says the constraint was unbounded variance.
Salesforce is reported to sell Agentforce per conversation at two dollars, as Flex Credits at 500 dollars per 100,000, and per user per month from 125 dollars, all concurrently. That is not a transitional state to be tidied later. It is what happens when a vendor cannot predict the cost of its own product well enough to pick one unit.
Red Hat, citing Intel, describes CPU-to-GPU ratios moving from one to eight in training toward one to one in agentic deployments, with some customers running four CPUs per GPU. Silicon Motion is positioning enterprise SSDs as a persistent tier for KV cache offload. Agentic cost is migrating out of the accelerator and into orchestration, memory and storage, which are budgeted by different people on a different cycle.
Cloudflare's GAAP operating margin moved from negative 13.1 percent to negative 29.6 percent on revenue of 696.1 million dollars, up 36 percent. A 150.7 million dollar restructuring charge accounts for roughly 21.6 points of that 16.5-point swing. Excluding it, operating margin improved year over year. The agentic build-out is visible in Cloudflare's product line, not in that margin line.
The seventeen-times token multiplier traces to an illustrative fraud-detection example, not a measured comparison against retrieval. The three-cents-to-fifteen-cents transaction cost has no locatable primary source. The 4.7x six-month increase attributed to Apptio in federal deployments originates in an April 2026 startup blog as a hypothetical. All three are reproduced with their audit rather than quietly dropped.
This report reproduces its own sourcing audit in full, because which numbers fail is itself part of the finding. Two of the three failures share a pattern: a plausible figure appears in a blog as an illustration, loses its hedging when quoted, acquires an institutional attribution it never had, and then circulates as a measured finding.
Traces to an illustrative fraud-detection example comparing an 800-token chat request with a 13,500-token agent task. That is agent versus chat, not agent versus retrieval. Published benchmark work with disclosed token counts puts agentic retrieval at 2.6x to 7.8x single-shot retrieval, and those are the most defensible figures available. They also happen to be the lowest, which is worth noticing.
Direction supported, magnitude unsourced, comparison misattributed
No primary source was located for either endpoint. The nearest sourced pair is EY's four cents in 2023 and one dollar twenty in 2026, which measures a different unit over a different period. The report substitutes the sourced pair and prints its limits rather than borrowing the credibility of a number nobody can find.
Replaced with the sourced pair, with its limits stated
No Apptio or IBM publication containing this figure was located. The number appears in an April 2026 startup blog as a hypothetical monthly bill growing from month six to month eighteen, presented as the author's own illustration. A hypothetical in a founder's blog post became a federal deployment statistic from an IBM company.
Dropped from the analysis entirely
The circulating figures were wrong in both directions. Output tokens moved from 0.87 dollars per million flat to 1.98 off-peak and 3.96 at peak, effective 16 August 2026. V4-Flash moved from 0.28 to 0.66 and 1.32. Corrected throughout, because the event matters more than the numbers and the numbers were being quoted as the event.
Corrected against the vendor's own rate card
Eight quarters out, with a subjective probability and an observable tripwire under each. The probabilities are the least reliable content in the report and are included because refusing to quantify a judgment is not the same as being careful. The tripwires are the part worth monitoring, and each can be watched without private information.
A platform fee sets a revenue floor at roughly seat-economics level, with agent consumption metered above it and customer-settable caps sold as a named feature. The dominant form because it keeps the vendor's forecastable base and moves the tail to the customer. Revenue recognition splits, a growing share of revenue stops being ratable, and remaining performance obligations stop covering forward revenue.
Tripwire: a top-ten application vendor discloses metered agent revenue as a separate line
Vendors hold per-seat pricing and cap agent behaviour at whatever depth the seat price supports, marketed as governance and predictability. Reported metrics keep working and gross margin is protected, but the product ships deliberately less autonomous than the demo. The comfortable scenario for public SaaS and the dangerous one for its five-year position.
Tripwire: pricing pages start advertising monthly action limits per seat rather than capability
Vendors charge per resolved ticket, reconciled invoice or closed candidate, and absorb the token variance themselves. The hardest accounting treatment of the four: consideration is variable and must be constrained under ASC 606, so recognition arrives later and lumpier. The vendor owns the entire right tail, including retries on tasks that never resolve.
Tripwire: a vendor reverses recognised variable consideration, or quietly redefines resolved
Sustained inference price increases across two or more providers with meaningful share, following DeepSeek. Not a spike but a durable reset of the floor as capacity binds. Multi-year fixed-price agreements signed against an assumption of falling input costs become loss-making, and the buyer-side consequence is a wave of price-increase notices framed as model cost pass-through.
Tripwire: a second provider with real share publishes peak and off-peak rates
The report also runs the gross margin arithmetic on a 40 dollar per seat per month agent product and puts the ceiling at roughly 33 completed tasks per seat per month before inference cost alone exhausts the subscription, under assumptions deliberately generous to the vendor. It carries a table of how each billing unit changes ASC 606 treatment, deferred revenue and gross margin behaviour; six leading indicators you can watch yourself in filings and rate cards; five stated assumptions with what fails if each is wrong; and four questions the published evidence cannot answer. The call is falsifiable and the condition is printed: I would abandon it if per-task token consumption flattens for two consecutive quarters while autonomous multi-step deployment keeps climbing.
This isn't a vendor summary. Every sentence is labeled by what stands behind it: verified fact, vendor claim, third-party estimate, my assessment, hypothesis, or scenario. Sources are numbered and clickable. Forward-looking sections use scenarios with observable tripwires, not forecasts. It's the same method behind every market assessment I write.
Twenty-four pages, built from public sources with no client brief and no interviews. Read it in the browser or take the PDF.
Each report here answers a real question, directed and researched against public sources and evaluated against a stated assumption, then delivered as Word and PDF. If you're weighing a platform, sizing a category, or defending a number to a board, tell me the decision behind it and I'll tell you honestly whether a report is the right tool.
Commission an assessment