Skip to main content
Robot hand struggles with complex multi-step task, symbolizing enterprise AI limitations and survey findings.

Editorial illustration for Survey: Most enterprise AI agents can't complete multi-step tasks independently

71% of Enterprise AI Agents Fail at Independent Tasks

Survey: Most enterprise AI agents can't complete multi-step tasks independently

4 min read

Seventy-one percent of enterprises told VentureBeat Research that a quarter or fewer of their deployed "agents" can actually finish multi-step work without a human stepping in. Only 10% said true agents make up most of what they run. These aren't casual guesses: 81% of respondents recommend or decide AI purchases at their companies, so they've seen the invoices and the failure logs.

The finding comes out of five parallel surveys VentureBeat Research ran in June, covering the five layers an enterprise needs before it can trust an agent with real work: identity, evaluation, cost telemetry, the context layer, and orchestration. Identity controls which agent can act and under whose credentials. Evaluation checks whether the output is any good.

Cost telemetry tracks what each agent run actually costs. The context layer feeds agents the business data they need to answer correctly. Orchestration ties multi-step agent work together.

Across every one of those five layers, enterprises are now scrambling to fix gaps they opened with their eyes open, and they're doing it with money: 57 to 68% plan to switch or add vendors within a year, with roughly a third moving inside the current quarter.

Enterprises deployed AI agents ahead of the controls needed to manage them — and they did it knowingly. That is the central finding across the five parallel surveys VentureBeat Research fielded in June, spanning every layer of the agentic stack.

Why this matters

The gap between what enterprises call an "agent" and what actually runs unattended is the real story here. Seventy-one percent of buyers admit a quarter or fewer of their deployments can finish multi-step work without a human stepping in, yet the label "agentic AI" got slapped on the budget line anyway. For builders, that's a warning about the tools you're shipping: procurement teams making these calls are senior enough to know better, which means the shortfall isn't ignorance, it's a governance problem that got greenlit anyway.

For founders selling into enterprise, the vendor-switching numbers (57 to 68% planning changes within 12 months) are the opportunity. Whoever can prove actual autonomy, not just chained API calls dressed up as agents, wins those contracts. Researchers should treat this as a data point on the distance between agent benchmarks and agent marketing.

The industry oversold autonomy before it built the controls to verify it, and now the retrofit budget is the tell. Watch which vendors get replaced first. That's where the definition of "agent" actually gets tested.

Common Questions Answered

What percentage of enterprise AI agents can complete multi-step tasks without human intervention?

According to VentureBeat Research, only 29% of enterprises reported that more than a quarter of their deployed AI agents can finish multi-step work independently. Seventy-one percent of respondents stated that a quarter or fewer of their agents could complete such tasks without human intervention, indicating a significant gap between expectations and actual autonomous capabilities.

Why did enterprises deploy AI agents before establishing proper governance controls?

VentureBeat Research found that enterprises knowingly deployed AI agents ahead of the controls needed to manage them, suggesting a disconnect between procurement decisions and operational readiness. The survey revealed that despite this gap, procurement teams making these calls are senior enough to understand the implications, indicating the shortfall stems from factors beyond ignorance rather than lack of awareness.

What does the survey reveal about the definition of 'agent' versus actual autonomous deployment?

The research highlights a significant gap between what enterprises label as an 'agent' and what actually runs unattended in production. Only 10% of respondents said true agents make up most of what they run, yet the label 'agentic AI' was still applied to budget lines, revealing a mismatch between marketing terminology and genuine autonomous capabilities.

How credible are the respondents in this VentureBeat survey regarding AI purchasing decisions?

Eighty-one percent of survey respondents either recommend or decide AI purchases at their companies, making them highly credible sources with direct visibility into invoices and failure logs. This high level of decision-making authority means their assessment of agent performance limitations reflects real operational experience rather than theoretical concerns.

LIVE22:23Box AI security concerns slow enterprise adoption of agentic tools