Somewhere in your stack right now there’s an agent pilot with no owner, no target metric, and a renewal date in Q3. That pilot is already dead; the invoice just hasn’t noticed. The numbers say this is the normal case, not the exception: McKinsey finds 62 percent of organizations experimenting with AI agents while 23 percent report scaling an agentic system anywhere in the enterprise, and in any single business function, fewer than 10 percent are scaling agents.

Your CFO will eventually ask which side of that funnel your spend is on. Better to ask it yourself first.

The funnel, with numbers on it

Walk the stages and watch the drop-off:

And the forward-looking number that should shape your contracts: Gartner predicts over 40 percent of agentic AI projects will be canceled by the end of 2027, citing escalating costs, unclear business value, and inadequate risk controls. S&P Global data points the same direction: the share of enterprises abandoning most of their AI initiatives jumped from 17 percent in 2024 to 42 percent in 2025, as reported in industry analyses.

The money is flowing anyway. Gartner forecasts agentic AI software spending up 141 percent in 2026 to nearly $202 billion. That combination, surging spend plus a 40 percent cancellation forecast, is not a contradiction. It’s a market paying tuition.

0%25%50%75%100%Using AI in ≥1 function88%Experimenting with agents62%Scaling agents anywhere23%Scaling in a given function10%
The agent pilot funnel, 2025-26 survey data Each value traces to the cited source in the text. The final stage is reported as under 10 percent. Source: McKinsey, The State of AI
The agent pilot funnel, 2025-26 survey data
CategoryShare of organizations
Using AI in ≥1 function 88%
Experimenting with agents 62%
Scaling agents anywhere 23%
Scaling in a given function 10%
Cite or embed this

Free to reuse with a credit link back to The Counter Brief.

Why pilots die: it’s the operating model, not the model

The post-mortems are remarkably consistent, and almost none of them blame the AI.

No owned metric. “Deploy an agent” is not an objective. The organizations getting value set outcome targets first. McKinsey’s high performers, the 5.5 percent of respondents attributing more than 5 percent of EBIT to AI, are the ones who tied AI work to a number someone owns.

No workflow redesign. Agents bolted onto an unchanged process automate the old process’s problems. High performers are 2.8x more likely to report fundamental workflow redesign, 55 percent vs 20 percent of others. That’s the single cleanest “do this” signal in the dataset.

Costs counted at the token, not the project. The pilot budget covers API calls. Production costs include integration, evaluation, monitoring, and the people reviewing outputs. Gartner’s cancellation drivers lead with escalating costs for a reason. We’re publishing a full cost breakdown of the agent stack on June 29; the short version is that tokens are the visible minority of the bill.

No kill criterion. Pilots without a pre-agreed failure condition don’t fail; they linger. Lingering pilots are the most expensive kind, because they consume the team’s credibility along with the budget.

The gate: five questions before any agent pilot starts

Run this before you sign, not at renewal. A “no” on any of the first four means don’t start yet.

  1. Which number moves? Name the metric (deflection rate, cycle time, cost per ticket, pipeline coverage) and the owner whose comp feels it.
  2. What’s the baseline? If you can’t measure the process today, you can’t prove the agent improved it. Instrument first.
  3. What changes about the workflow? If the answer is “nothing, the agent just helps,” you’re buying a 23-percent-survival-rate lottery ticket.
  4. What kills it? A date and a threshold. “If audited accuracy is under X by week 8, we stop.”
  5. What does production cost? Not the pilot. Production: integration, eval suite, review labor, monitoring. Ranges are fine; silence is not.

Where this framework breaks, and you should know its edges: it under-serves genuine R&D. If you’re deliberately exploring capability with no near-term metric, exempt that work explicitly, cap it as a research budget, and don’t let it masquerade as an ROI pilot. The framework also can’t rescue a pilot built on a vendor that isn’t real; Gartner estimates only around 130 of thousands of self-described agentic vendors actually are. Vendor diligence is a separate exercise (our June 20 piece covers it).

What to do Monday

If you run a team under ~200 people: inventory your agent pilots. For each, write the metric, owner, baseline, and kill date on one page. Any pilot where you can’t fill all four fields gets 30 days to earn them or it’s cut. You likely have one or two pilots; this takes an afternoon.

If you run revenue ops at scale: same inventory, plus two additions. Move every agent contract you can to outcome-aligned or short-cycle terms ahead of the renewal wave, and pick one workflow (not one tool) for genuine redesign in H2. The 2.8x redesign signal is the highest-value fact in this piece; resource it like you believe it.

Either way, the position to be in by Q4 is simple: every agent dollar in your budget maps to a metric, an owner, and a kill switch. The 40 percent that get canceled will mostly be the dollars that didn’t.

FAQ

What percentage of AI agent pilots reach production? McKinsey’s State of AI survey shows 62 percent of organizations experimenting with AI agents while 23 percent report scaling an agentic system anywhere in the enterprise, and fewer than 10 percent are scaling agents within any single function. Separately, Gartner predicts over 40 percent of agentic AI projects will be canceled by the end of 2027. Roughly one in four experimenters reaching scale is the honest planning number.

Why do most AI agent projects fail? The dominant causes are organizational, not technical: no owned business metric, no baseline measurement, no workflow redesign, and costs that escalate beyond the pilot budget. Gartner cites escalating costs, unclear business value, and inadequate risk controls as the drivers behind its 40 percent cancellation prediction. McKinsey’s high performers are 2.8x more likely to have fundamentally redesigned workflows around AI rather than bolting agents onto existing processes.

How do I measure ROI on an AI agent? Define one target metric and its owner before the pilot starts, measure the baseline first, and set a kill threshold with a date. Count full production costs (integration, evaluation, human review, and monitoring, and not tokens alone) against the metric’s movement. If you cannot name the metric, the baseline, and the kill criterion on one page, the pilot is not measurable and should not start.

The Counter Brief — one email, every Monday.

The week's AI-for-revenue moves in a 5-minute read: which tools are worth the budget and which to skip, plus what to do this week. Source-checked, no vendor decks.

Edited by Aditya Marin Gasga

Free. One click to unsubscribe.

About Nishtha Gupta

Contributor · Operations Lead, Demand Nexus

Nishtha Gupta leads an operations team at Demand Nexus, running appointment generation and day-to-day execution against targets. Her background is in B2B data operations and analysis, including over a year as a B2B research analyst at Ziff Davis Performance Marketing. She is completing a master's in computer applications at MIT World Peace University.

More from Nishtha Gupta →