Salesforce counts tasks. Workday counts customers with one thing switched on. Microsoft counts registry entries. ServiceNow counts contract value tagged to a bundle the customer did not separately buy. Where a platform publishes both what was built and what got reused, reuse is 2.5 percent of creation.
Salesforce counts discrete tasks executed in production. Workday counts customers with at least one agent switched on. Microsoft counts registry entries. ServiceNow and Workday both count contract value that somebody inside the vendor tagged as AI. Synapxe counts agents built by named professionals and the subset published for anyone else to use. Those are five different objects with five different floors, and they get printed next to each other in board decks and sell-side notes as though they were the same measurement.
The definitional layer never arrived. Not one of these metrics appears in a filed periodic report. Salesforce's “Agentic Work Unit” returns three hits in the SEC EDGAR full-text index, all of them Exhibit 99.1 earnings releases attached to a Form 8-K, and the cover of that 8-K states the information “shall not be deemed 'filed' for purposes of Section 18.” Workday's “organic agents” sits in an 8-K exhibit too. The usual comfort that a number has been through the disclosure controls of a 10-Q does not apply to any figure in this study.
What makes the gap hard to see is that these metrics have a floor and no ceiling. A count of customers using one or more agents cannot fall when a vendor retires agents or merges them, which is exactly what Workday's chief executive says happened. A registry count cannot fall when an agent is created, tested once and abandoned. Six of the eight metrics graded here score one or zero out of four on definition, denominator, stability and checkability. That is not a category where cross-vendor benchmarking is possible, however confidently the comparison gets drawn.
Built from earnings releases, call transcripts, the SEC EDGAR full-text index and platform operator announcements covering Salesforce, Workday, ServiceNow, Microsoft, Okta, nCino, IFS and Synapxe. Each finding is labeled by evidence type, so a figure verified in a filing is graded separately from a vendor's characterisation of its own business and separately again from my reading of the two together. Where the evidence cuts against the argument, it is printed rather than dropped.
A full-text search of the SEC EDGAR index returns Salesforce's term “Agentic Work Unit” in exactly three documents, all Exhibit 99.1 earnings releases attached to Form 8-K. “Agentforce and Data 360 annual recurring revenue” returns six. Workday's “organic agents” appears in its Q2 8-K exhibit. None appears in a 10-Q or a 10-K, and the Salesforce 8-K cover disclaims Section 18 liability for the contents.
Synapxe, Singapore's national healthtech agency, reported more than 12,000 AI agents created on its AgentSea platform between the end of May and late August 2026, and more than 300 of those made available to other users for reuse. The eligible population is roughly 80,000 public healthcare professionals. It is the only disclosure in the study that publishes a numerator, a reuse figure and a population together.
Satya Nadella told the FY26 Q4 call that “nearly 40 million agents registered across tens of thousands of companies within two months of launch.” Reading “tens of thousands” at a midpoint of 50,000 companies implies about 800 registered agents per company. Salesforce's own production telemetry puts activated agents at 13 per business as of April 2026. Registration and activation are not the same event, and Microsoft does not say which it counts.
The Q2 FY27 release states: “Beginning in Q2 FY27, Agentforce ARR includes our AI offerings Slackbot and Headless 360.” The same release reports Agentforce ARR above $1.5 billion, up over 240 percent year over year. A growth rate spanning a composition change is not a like-for-like comparison, and the release does not quantify what the newly included products contributed.
On 9 April 2026 ServiceNow replaced five legacy tiers with three, and Now Assist, the Moveworks layer, Workflow Data Fabric and AI Control Tower were bundled into every tier rather than sold as add-ons. Once AI ships inside the base SKU, contract value tagged to AI reflects a pricing decision by the vendor rather than a purchasing decision by the customer. The $1 billion AI ACV milestone spans that change.
Aneel Bhusri told diginomica: “We had a lot more agents when I came back. We killed a lot or we rolled them into bigger agents.” Workday's published metric is customers using one or more agents, which is structurally incapable of falling when agents are retired or merged. Neither the starting agent count nor the number removed was disclosed.
Harvard Business Review Analytic Services, surveying 603 leaders with sponsorship from Workato and AWS, found 6 percent fully trust AI agents to run core business processes autonomously, 43 percent limiting them to routine tasks and 39 percent to supervised or noncore use. Futurum research commissioned by IFS put the equivalent figure at 5.7 percent. Both sponsors sell agent software, which makes the low number harder to dismiss rather than easier.
Okta's Q2 commentary described “several AI-related deals in the quarter with contract values above the corporate average” without an AI-specific ARR, customer count or transaction volume. nCino reported four US enterprise customers holding over $900 billion in combined assets renewing early with “expanded AI commitments” and quantified none of it. Neither published a number that could be misread, because neither published a number.
Where a definition exists it is quoted as the vendor printed it. Where none exists, the absence is stated rather than filled in. The grading rubric has four criteria, each worth a point: is a definition published, is a denominator published alongside the numerator, did the composition stay stable across periods, and can an outsider check the figure against another disclosure. Six of the eight scored metrics come in at one or zero.
Published verbatim as “a measure of discrete tasks executed by AI agents in production across the Salesforce platform, including Agentforce and Slack.” The best definition in the set, and still silent on what a unit is worth: resolving a customer case and updating a record are two or three orders of magnitude apart in value, and both count as one.
Denominator: none. No customer count, no tasks-per-customer split, no pre-agent baseline
A blended definition covering Data 360 plus “certain generative AI subscription agreements,” with an Agentforce-only carve-out given separately. The metric carrying the written definition is the blended one; the metric quoted in headlines is the carve-out, and the carve-out absorbed Slackbot and Headless 360 mid-year.
Denominator: none. Composition changed in Q2 FY27 and the added products were not quantified
No definition. “Organic” is load-bearing, appears to exclude agents from acquisitions and partners, and is nowhere explained in the release. A threshold of one is close to a measure of feature availability: it rises when a customer switches something on and cannot fall when they stop.
Denominator: not published. Roughly 11,000 customers, so about half the base, by the reader's arithmetic
Neither “AI” nor “new ACV” is defined in the release. The tagging happens inside the vendor and is never seen by the customer, so when a customer buys a suite that includes agents, the split between the AI line and the base line is an internal decision with no external check.
Denominator: new ACV in the period, itself undisclosed
No definition located in the earnings materials. The metric spans the 9 April 2026 packaging change that bundled Now Assist into every tier, so it mixes a period when buying AI was a customer decision with a period when it was not. Legacy SKUs reached end of sale on 1 July 2026.
Denominator: total ACV, not published alongside
Neither “agent” nor “registered” is defined on the call. A registry entry is the cheapest possible unit: no execution, no user, no business process. It is the correct thing for Microsoft to track internally, because governing agents is the product, and the wrong thing to place next to a competitor's production or revenue figure.
Denominator: “tens of thousands of companies,” a range spanning a factor of five
Implicit but unambiguous: agents built by professionals, and the subset published for others to use and adapt. Sharing is a deliberate act, so the 300 understates agents used privately and never published. It is still the cheapest proxy available for an agent being good enough that a second person should have it.
Denominator: roughly 80,000 eligible public healthcare professionals, published
Digital workers are defined as agents that “autonomously monitor, decide, and execute operational tasks, escalating to humans only when judgment is required.” A genuinely useful disclosure shape, undermined by the missing transaction base: the ratio can hold steady while the volume behind it moves in either direction.
Denominator: agentic transactions on the IFS platform, count not published
Okta references “several AI-related deals” above average contract value. nCino reports four renewals with “expanded AI commitments.” Neither is a metric, and both companies are excluded from the scoring because there is nothing to grade. Their absence from the chart is not a failing grade, and it is the more defensible position.
Denominator: not applicable
Point forecasts are not appropriate here, because the outcome turns on a small number of discretionary decisions by vendors and regulators rather than on a trend that can be extrapolated. The probabilities are subjective and are the least reliable content on the page. The earliest visible signs are the part worth monitoring, and each is observable from published documents within days of release.
As year-over-year growth rates decelerate, vendors stop reporting agent counts and shift the narrative to outcome metrics such as cases resolved or cost per resolution, without ever restating or reconciling the counts they published in 2026. Comparability is lost permanently. Anyone who built a peer benchmark on 2026 agent counts finds the series ends mid-air, and the 2026 figures become uncheckable against anything.
Sign: a vendor that reported an agent metric for three consecutive quarters omits it from a release while still discussing agents in the prepared remarks
Pressure from auditors, SEC comment letters on Item 2.02 metrics, or a major analyst firm publishing a standard pushes vendors to footnote agent metrics the way they footnote ARR, with stated denominators and disclosure of composition changes. Cross-vendor comparison becomes possible for the first time, and several 2026 growth rates look materially worse when restated on a consistent basis.
Sign: any agentic operating metric appearing inside a 10-Q rather than only in a furnished 8-K exhibit, or an SEC comment letter naming an agent metric
Registry-style counts become the category standard because they produce the largest numbers, until a short-seller report, a restatement or a customer-side audit publicly contradicts a headline figure and forces a vendor to explain what it was counting. A sharp repricing of AI attach-rate assumptions across the sector follows, and with it a rapid, involuntary move to the disclosure discipline of scenario B.
Sign: a published bottom-up reconstruction of a vendor's agent count from customer interviews landing materially below the disclosed figure
The report also carries a scored disclosure-quality chart across eight metrics, seven leading indicators with what each one means and where to watch it, and a reconciliation of the adoption numbers against the trust surveys. It sets out separate implications for CIOs benchmarking peer adoption, equity analysts pricing AI attach rates, procurement leads negotiating agentic renewals, product leaders setting adoption targets and architects scoping pilots. It states four questions it cannot answer without platform telemetry or vendor interviews, prints a half-life on every figure in it, and names the weakest assumption in the document rather than burying it. The call is falsifiable and the condition is printed: what would move me is a vendor footnoting its agent metric in a filing and holding that definition through a quarter of decelerating growth.
This isn't a vendor summary. Every sentence is labeled by what stands behind it: verified fact, vendor claim, third-party estimate, my assessment, hypothesis, or scenario. Sources are numbered and clickable. Forward-looking sections use scenarios with observable tripwires, not forecasts. It's the same method behind every market assessment I write.
Twenty-two pages, built from public sources with no client brief and no interviews. Read it in the browser or take the PDF.
Each report here answers a real question, directed and researched against public sources and evaluated against a stated assumption, then delivered as Word and PDF. If you're weighing a platform, sizing a category, or defending a number to a board, tell me the decision behind it and I'll tell you honestly whether a report is the right tool.
Commission an assessment