Every major AI lab and most enterprise software vendors now have an agent story. Not a chatbot - an agent. Something that takes a goal, breaks it into steps, calls tools, browses the web, writes and executes code, sends emails, and reports back when it’s done. The pitch is efficiency. The reality being quietly papered over is that nobody has a reliable way to audit what these things actually do between the prompt and the result.

This isn’t a theoretical concern. Salesforce’s Agentforce, Microsoft’s Copilot agents, and a widening field of startups are already running in production environments where they have access to CRMs, inboxes, calendars, and internal databases. The agents act on behalf of employees. They can initiate actions - not just answer questions. And most of the organisations deploying them are evaluating them on task completion, not on what decisions were made along the way.

The logging problem is underappreciated. A language model that answers a question leaves a record: the prompt, the response, maybe a token count. An agent that completes a multi-step task generates a chain of internal reasoning, tool calls, intermediate outputs, and decisions that most current deployment frameworks don’t surface to the operator in any usable form. You see the outcome. The path is largely opaque.

This matters more than it sounds because agent errors compound. A model that gives a bad answer is annoying. An agent that misinterprets a goal and then takes twelve sequential actions based on that misinterpretation can cause damage that takes hours to untangle - deleted records, sent emails, modified entries in systems of record. Several companies that adopted early agentic workflows have reported exactly this kind of cascading failure, though most haven’t done so publicly.

The speed of deployment is being driven partly by competitive pressure. If your competitor is claiming their AI handles entire workflows autonomously, you’re under pressure to say the same - whether or not your observability tooling has caught up. It mostly hasn’t.

What’s missing isn’t better models. The capability is largely there. What’s missing is the operational layer: structured audit trails, meaningful human-in-the-loop checkpoints for high-stakes actions, and rollback mechanisms that actually work. A handful of smaller tools - LangSmith, Weights & Biases’ agent tracing, some internal enterprise builds - are working on this. But they’re far behind the deployment curve.

The industry built the plane while it was already in the air. Now it’s discovering the instruments are missing.