Most boards now hear that their competitors are “doing AI”. Fewer hear a harder fact: scaling AI agents beyond a tidy pilot is still uncommon. Headlines blur adoption with impact. The numbers do not.
McKinsey’s State of AI survey for 2025 finds that 88% of organisations use AI in at least one function, yet only 23% are scaling an agentic AI system somewhere in the enterprise. An additional 39% are experimenting. In any given function, no more than 10% report scaling agents. That gap is the story operators need to read carefully.
What scaling AI agents actually means in the data
McKinsey treats agents as systems based on foundation models that can plan and execute multiple steps in a workflow. “Scaling” here is not a press release. It means expanding deployment and adoption within at least one business function. “Experimenting” is earlier: teams are trying agents without that broader rollout.
Put the layers together and the picture sharpens. Sixty-two percent of respondents say their organisations are at least experimenting with AI agents. That curiosity figure is the sum of those scaling somewhere and those still in trials. It is not a claim that agents already run the business. Curiosity is widespread. Production depth is not.
The function-level ceiling matters even more for founders and SME owners. A firm can honestly say it is “scaling agents” because one workflow in IT or knowledge management is expanding, while every other function sits at or below that 10% band. Enterprise-wide agent operations remain rare. Local progress can look like a transformation story when it is really a single corridor of light.
Why the 88% figure misleads boards
Regular AI use in at least one function is now the norm in McKinsey’s sample. That includes chat assistants, content tools, coding helpers, and analytics copilots that never leave a named team. Those tools create real time savings. They do not automatically create multi-step agents that write back to systems of record, trigger actions, or own a workflow end to end.
Boards that treat the 88% as proof they are behind on agents confuse two products. Assistive AI raises individual productivity. Agentic systems change who does which steps, and under what controls. The first is easy to pilot with a licence. The second needs ownership, data access, kill switches, and a metric that survives live customers.
That distinction also explains why investment narratives and operating reality diverge. Capital still follows credible operating discipline more reliably than demos, a theme we explored in our look at AI investment trends for 2026. A crowded demo day does not change the McKinsey ratios. It only raises the noise around them.
Reading 23%, 39%, and the 10% ceiling together
The 23% scaling figure answers a yes or no question: is an agentic system expanding in at least one function? The 39% experimenting figure describes parallel trials that have not crossed that bar. Together they form the 62% “at least experimenting” headline. Separately, they show that most agent work is still early.
The at most 10% ceiling per function then cuts through vanity metrics. If scaling is concentrated in one or two functions, the rest of the company may still run on ticket queues, shared inboxes, and heroics. For a mid-sized firm, that pattern is familiar. One team ships a working agent for service-desk triage or research summaries. Neighbouring teams keep the same manual handoffs they had last year.
McKinsey also notes that agent use shows up most often in IT and knowledge management, with wider reporting in technology, media and telecommunications, and healthcare. That concentration should temper copy-paste playbooks. A knowledge-management research agent does not prove your quote-to-cash process is ready for autonomous steps.
What stalls the move from pilot to production
Independent surveys tell the same story from different samples. Deloitte’s Tech Trends 2026 agentic AI brief reports that while 38% of organisations are piloting agentic solutions, only 11% are actively using these systems in production. Gartner separately forecasts that task-specific AI agents will sit in 40% of enterprise applications by end-2026, up from less than 5% in 2025. Those are not McKinsey’s ratios restated. They are a second, independent confirmation that pilots are common and production depth remains thin.
Pilots fail for boring reasons. There is no named process owner. The agent cannot write to the CRM, ERP, or case system. Nobody defined a kill criterion, so a weak trial lingers. Staff do not trust the underlying data, so every automated suggestion gets re-checked by hand until the time saving disappears.
Those failure modes are operational, not model-centric. Model quality still matters, but it is rarely the binding constraint once a pilot has produced a plausible demo. The binding constraint is whether the firm redesigns the workflow the agent is supposed to run inside. We covered that production gap in detail in moving agentic AI from pilot to production, and the McKinsey ratios are the quantitative backdrop for that argument.
Data foundations decide how far an agent can go without creating new mess. If customer records, inventory states, or ticket histories are incomplete, an agent that acts on them will accelerate error. Building cleaner first-party systems is therefore not a side project for marketing analytics alone. It is a prerequisite for safe automation, which is why first-party data systems that drive B2B growth sit next to any serious agent roadmap.
A practical reading for operators this quarter
Start by separating three claims in any vendor or internal update: AI use, agent experiments, and agent scaling in a named function. Ask which McKinsey layer the claim maps to. If the answer is “we use AI”, you are in the 88% band. If the answer is “we are trying agents”, you are closer to the experimentation slice. Only a measurable expansion of deployment in a live workflow belongs in the scaling bucket.
Next, pick one process with volume, cost, and a single owner. Write the baseline cycle time, error rate, and rework hours. Decide which steps may never be automated without a human approval, especially money movement, legal commitments, and irreversible customer actions. Only then attach an agentic layer to the remaining steps.
Finally, refuse enterprise-wide agent programmes until one function shows a metric that moved and stayed moved. The at most 10% finding is a reason to prefer depth over theatre. One scaled workflow with an owner, a system of record, and an exit rule beats five disconnected pilots.
What the numbers do not say
They do not say agents are a fad. Curiosity is high, and high performers in McKinsey’s framing are far more likely than peers to report scaling agents in most functions. They do not say SMEs should wait for perfect governance frameworks before starting. They say start small, instrument honestly, and do not confuse a licence rollout with operating change.
They also do not license invented urgency. Most organisations use AI somewhere. A minority scale agents somewhere. Almost none scale them broadly across functions. That is the sober baseline for capital, headcount, and risk.
For investors and profitable SME owners, the useful question is not “are we behind on AI?” It is “which workflow will we scale an agent inside, on what metric, and under whose authority?” Answer that, and the McKinsey chart becomes a filter for where effort belongs.
Key takeaways
- Scaling AI agents is not the same as using AI. McKinsey reports 88% AI use in at least one function against 23% scaling an agentic system somewhere.
- Sixty-two percent are at least experimenting with agents (23% scaling plus 39% experimenting). Curiosity is common. Depth is not.
- In any given function, no more than 10% report scaling agents, so “we are scaling” often means one corridor, not the whole firm.
- Pilots stall on ownership, data access, and kill criteria more often than on model choice.
- Operators should map every claim to use, experiment, or scale, then deepen one measurable workflow before buying a company-wide programme.
FAQ
What does McKinsey mean by scaling AI agents?
Scaling means expanding deployment and adoption of an agentic AI system within at least one business function, not a one-off pilot or a team demo.
How can 62% experiment with agents if only 23% are scaling?
The 62% figure is those at least experimenting, which combines the 23% already scaling somewhere with the additional 39% still in earlier trials.
Why does the 10% function-level figure matter more than the 23%?
The 23% can reflect progress in one or two functions. The at most 10% ceiling shows how rare scaling remains inside any single function across the survey base.
Does high AI use mean a firm is ready for agents?
No. Assistive AI use and multi-step agents are different layers. Ready firms have a named owner, clean data access, and a kill switch for the target workflow.
Where should a mid-sized firm start with agent scaling?
Choose one costly, high-volume process with a single owner, fix the system of record, set a success metric and exit rule, then expand only after the metric moves.
Sources
- McKinsey, The state of AI in 2025: Agents, innovation, and transformation, 5 November 2025
- Deloitte Insights, Agentic AI strategy (Tech Trends 2026), 2025
- Gartner, 40% of enterprise apps will feature task-specific AI agents by 2026, 26 August 2025
This article is general information, not investment advice. See section 4 of our Terms & Conditions.
This is the kind of decision we advise on. If you are weighing it for your own business, see how we work or start a conversation.
