A number went around the internet last week: 49 percent of executives have scaled back AI agent deployments because operating costs outweighed the benefits. It came from KPMG's Global AI Pulse for Q2 2026, a survey of 2,145 senior leaders across 20 countries, and within a day Polymarket was quoting odds on an AI bubble burst.
Sandy Carter read the full report so the rest of us didn't have to, and her Forbes breakdown is the one to read. Her line is the correct frame: "That is not a bubble bursting. That is a bill arriving."
She is right. Look at what else sits in the same dataset. AI is a top investment priority for 79 percent of leaders, up from 74 percent the quarter before. Average AI budgets are holding at $188 million. The share of organizations calling AI part of everyday work jumped from 13 percent to 22 percent in a single quarter, the biggest move KPMG has ever measured at any point on its maturity curve. Nobody is leaving. They are rephasing.
We have seen this movie. A decade ago, cloud invoices arrived that nobody could explain, and the industry answered with FinOps. Owners, forecasts, unit costs, a discipline. The token era is getting its own version, and KPMG's data shows it forming in real time. 53 percent of organizations report AI cost dashboards. 54 percent have embedded cost review in their AI approval loops. The companies with dashboards are roughly five times more likely to report real ROI.
Here is where I part ways with the standard prescription.
Phase one and phase two
Cloud FinOps matured in two distinct phases. Phase one was visibility. See the spend, tag the resources, show back the costs. Phase two was where the money actually got saved: rightsizing, reserved instances, deleting the zombie infrastructure that phase one exposed. Visibility never saved a dollar by itself. It told you where the dollars were dying.
Every recommendation now circulating for AI cost, including the four in the KPMG coverage, is a phase one move. Install a meter before you scale. Make token economics a leadership literacy. Put cost review inside the approval loop. Rephase rather than retreat. All correct. All observability.
Nobody is talking about phase two, because phase two requires knowing what the meter will reveal once you can finally read it. And what it reveals is uncomfortable.
What the meter shows
Agents are priced like electricity but they do not behave like appliances. They run long tasks, call other tools, and check their own work, and every step is on the meter. When GitHub Copilot moved to usage-based billing on June 1, a Visual Studio Magazine writer tracked his first day and projected a $180 monthly bill on a plan that had been flat $10. One long, tool-heavy session. Eighteen times the cost.
Now ask what those metered steps actually were. Some fraction was genuinely new work, novel reasoning on a novel problem. Look inside any agent deployment, though, and a large share is repeat purchases. The agent re-derives a policy it derived yesterday. It re-verifies a fact the organization verified a thousand times. It re-retrieves and re-summarizes the same document for the fortieth user this month. Each of those steps bills full freight and produces zero new value.
This is the Rediscovery Tax, and it is invisible on a cost dashboard. The dashboard tells you the bill is high. It does not tell you that you bought the same answer 400 times.
That is why only 26 percent of billion-dollar companies reporting full real-time cost visibility is not even the scary number. The scary number is how few of the 26 percent can decompose their spend into first purchases versus repeat purchases. Token Yield, the ratio of necessary spend to total spend, is the metric phase two runs on. Almost nobody can compute it yet.
What phase two looks like
Phase two of cloud FinOps was rightsizing and reservation. Stop paying on-demand prices for predictable workloads. Phase two of token FinOps is interception. Stop paying inference prices for answers you already own.
The mechanics differ from cloud because the waste differs. Cloud waste was idle capacity, machines running with nobody using them. Token waste is rediscovery, intelligence re-manufacturing its own prior output. You fix idle capacity by turning things off. You fix rediscovery by remembering, which means a memory layer that sits in front of the model, recognizes when a question has already been answered, and resolves it from knowledge the enterprise owns instead of renting the answer again.
This is also where cost discipline and governance stop being separate conversations. An answer resolved from your own knowledge graph never left your perimeter, never touched a vendor's meter, and never depended on a model's mood that day. Rent the commodity. Own the differentiation.
The bill is the beginning
KPMG's data does not describe a retreat. It describes a market that just learned to read its first invoice, and cloud already taught us the discipline takes two phases to complete.
The 49 percent who pulled back agents are the realists in this dataset, clearing room to scale what pencils out. When they come back, and the 79 percent priority number says they will, they will arrive with a phase two question. Not "what did the agents cost," but "how much of that did we need to spend at all."
That question has an answer, and it is a lot smaller than the current bill.
The 49/79/26 figures are from KPMG's Global AI Pulse Q2 2026, surfaced in Sandy Carter's Forbes analysis, which includes a plain-English token pricing explainer your CFO will thank you for.

