Home 9 AI 9 When AI Agents Fall Short of Expectations

When AI Agents Fall Short of Expectations

by | Jan 28, 2026

Why mathematical limits and hallucinations challenge the promise of autonomous AI systems.
Source: Wired Staff; Getty Images.

 

In a recent Wired.com article, veteran tech journalist Steven Levy looks at the gap between grand industry claims about AI agents and the current reality of their capabilities. Major companies marketed 2025 as the year when autonomous AI assistants would take on real-world work, but the technology has not delivered on that promise at scale. One reason is fundamental: a new research paper argues that transformer-based large language models are mathematically unable to reliably carry out complex autonomous tasks. The authors of that work suggest that no matter how you enhance reasoning in these systems, they cannot be made fully dependable for arbitrary agentic behavior under all conditions.

The article highlights this tension by quoting Vishal Sikka, a former tech executive and co-author of the paper, who says such AI simply cannot be fully reliable, making them unsuitable for safety-critical operations. At the same time, the industry pushes back. Tech leaders and startups point to progress in reducing hallucinations and in specialized agent applications, especially coding tasks where structured verification and formal methods can improve trustworthiness.

One startup, Harmonic, uses formal verification via mathematical languages to make outputs provably correct in specific contexts, a step toward greater reliability. Even so, hallucinations, incorrect but confident output, remain a persistent challenge across models, and internal research from major AI labs confirms that perfect accuracy is unreachable.

Levy’s article underscores that neither outright rejection nor blind optimism captures the state of agentic AI. Instead, progress is uneven: agents are becoming more capable in narrow domains, but fundamental limits mean widespread, dependable autonomy is still out of reach. Hallucinations and verification overhead erode practical value, so many businesses remain cautious about broad deployments.

The article closes with a broader reflection: rather than seeking a single breakthrough year, observers may need to think about how AI agents reshape work and cognition over time, even if they never fully match human reliability.