Home 9 AI 9 AI-Run Businesses Expose the Limits of Autonomous Agents

AI-Run Businesses Expose the Limits of Autonomous Agents

by | Sep 15, 2026

Andon Labs puts AI agents in charge of stores and cafés to study the gap between completing individual tasks and operating reliably in the unpredictable real world.
Andon Café in Stockholm employs human workers, but an AI agent named Mona manages the budget, orders supplies, and sets the menu (source: Andon Labs).

 

Andon Labs is testing autonomous AI in an unusual setting: real businesses. The San Francisco AI safety company operates stores and cafés where AI agents manage activities such as inventory, pricing, menus, budgets, deliveries, and vendor communications. These businesses serve as experiments for understanding how much responsibility current AI agents can handle without constant human supervision, tells IEEE Spectrum.

The work grew from Vending-Bench, a 2025 simulation in which agents based on models from Anthropic, Google, and OpenAI managed virtual vending-machine businesses. Researchers found that performance often deteriorated over time. Agents forgot orders, misunderstood delivery schedules, entered repetitive failure loops, and sometimes rationalized questionable behavior.

Andon moved into physical businesses because real environments introduce situations that simulations cannot easily anticipate. At Andon Market in San Francisco, an AI manager named Luna tracks deliveries and communicates with vendors while human employees perform physical tasks. Luna has demonstrated useful flexibility but also makes basic perception mistakes, such as repeatedly identifying an electrical floor cover as a loose coaster.

Andon Café in Stockholm has revealed other limitations. A Gemini-based manager initially bought excessive fresh ingredients that spoiled. After switching to a GPT-based agent, the system overcorrected, avoiding perishable ingredients and reducing the menu largely to cheese toast.

These experiments highlight a critical distinction between AI capability and reliability. An agent may successfully perform individual tasks yet still struggle to manage an organization consistently over time. Princeton researcher Sayash Kapoor argues that reliability is improving more slowly than capability, making it an important metric for evaluating autonomous systems.

Andon plans to use problems discovered in its physical businesses to create digital twins where failures can be reproduced under controlled conditions. The combination could help researchers identify unexpected behaviors in the real world, then systematically test whether new models and safeguards make autonomous agents more dependable.