AI Agents Have a Lemons Problem

Written by Dr. Andre Wenz | 7 min read
Last modified: September 1st, 2026
Smarter Processes for the Autonomous Enterprise: AI Capabilities and Agents in SAP Signavio Americas/EMEA header image

No one checks whether their toaster is wired safely. We trust that it will not burn down our kitchen. That trust took fires to earn.

In 1893, Chicago built a palace to celebrate electricity, with over 150,000 light bulbs on display. This was the largest showcase of electricity assembled at that time. Less than a year later, it burnt down.

No appliance manufacturer wants their products to burn down houses, so why were they so unsafe? Because safety was invisible at the point of sale. A carefully wired appliance and a dangerous one looked identical in the shop and on the price tag. The manufacturer who spent on safety priced himself out of the market. The one who skipped it won the sale. Quality you can't see is quality you can't charge for.

Cheap Cognition Makes Quality Harder to See

History might not repeat itself, but incentives do. Today, that wiring is now cognition, and the shop floor is your process landscape. Every company is in a mad dash to put AI agents into production. The next frontier of efficiency is bending tokenized cognition into the shape of your business workflows. Yet again, we have few ways to tell which agents are reliable and which are a process firestorm waiting to happen.

Gartner expects more than 40% of agentic AI projects to be canceled by the end of 2027, citing costs, unclear business value, and inadequate risk controls. But agents are not appliances. They are labor. When competence is free to fake, trust becomes expensive.

No One Is Coming to Certify Your Agents

The economic theory behind this dilemma is known as the lemons market and was awarded the Nobel Prize in 2001. When buyers cannot tell quality apart, everything sells at the average, and the maker of the good version cannot recover their costs. He cuts quality, or he leaves.

The same dynamic now plays out inside the company. An agent built with deep context, evaluation, monitoring, and human escalation costs more than a thin wrapper around a model. If management cannot see the difference in outcomes, the better-engineered system appears inefficient and gets cut first. As cheaper AI agents scale, average quality falls. Worst of all, markets that reach this state tend to remain there.

So, who supplies the missing signal? Two options, neither viable. Waiting for universal standards means accepting the cost of delay. Trusting the providers without independent evidence means accepting the cost of failure. Regulators will not solve the problem in time. Standards move at the speed of law, whereas agents move at the speed of a release cycle. Even when the rules arrive, they can set a floor at best. They cannot tell you whether an agent is good at your company's specific work.

Asking AI providers fails for obvious reasons. Asking a vendor to certify its own agent's reliability is like asking a used-car dealer to write an inspection report. Besides, their benchmarks measure performance under their conditions, not yours.

This leaves only one door open. Companies must build their own evidence while they deploy. The signal has to come from you.

Reputation Is Not Operational Proof

Most companies don’t have the luxury of waiting for a fix to arrive. The risk of innovation eroding their business model is too acute. So rather than pay for unfounded trust, many buyers overpay at the high end of the market. Lacking clear quality signals and objective performance, they stick to brand promises.

Frontier AI labs are known for selling the public on dystopian risk scenarios, signaling that they care more than anyone else, even when they don't provide effective risk measures. The brand comes to stand for safety, and in a world without clear measures, feelings translate neatly into price premiums.

The true danger lies in the gap between label and reality. It can become cheaper to invest in the brand than to build what the brand promises.

Stop Buying Trust, Start Measuring It

What breaks the cycle? Not another vendor questionnaire. Companies need to observe what an agent does, compare it with the intended outcome, calculate the cost of its failures, intervene before an error becomes a pattern, and turn every correction into something the next agent can build upon.

The relevant unit of governance is not the model. It is the end-to-end execution of the process with the full business context. Otherwise, businesses might focus on the wrong signals, such as 10x speed-ups in agentic invoice creation, without knowing whether the long-term failure rate turns the expected win into a verifiable loss. Model quality matters, but operational quality is what the company owns.

Here is the counterintuitive part. Replacing human labor with agentic labor should raise quality. But if you cannot evaluate that labor, the opposite happens. A company that moves from human labor to agentic labor without investing in telemetry does not simply get the same work for less. It loses the ability to tell good work from plausible work. Management begins pricing and optimizing for the average, the quality premium disappears, and the company’s standard of work gradually falls. This is why observability is not a support function for agentic AI. It is the institution that prevents an internal lemon market. When reliable performance is observed, teams can justify investing in context, controls, and evaluation. When it cannot, those investments can look like overhead and be stripped out. Measurement does not merely reveal quality. It preserves the incentive to raise it.

Memory Is the Control Layer

This is why we invest in SAP Company Memory. Company Memory turns institutional knowledge into an execution asset. It captures the rules, preferences, exceptions, escalation logic, and trade-offs that define how the company wants work to be done. Process Atoms make that knowledge small enough to govern and precise enough for agents to use.

A quality signal needs something stable to measure against, and in an enterprise, that something is the company's own operating knowledge. Not just the facts, suppliers, contracts, and policies an agent can already look up, but the procedural layer underneath: how decisions are made, when to escalate, and what a correct outcome actually is. Company Memory captures that layer so that an agent's work can be graded against how the business truly runs.

It improves with use. Every agent action leaves a trace, and the most valuable failures expose exactly where the company's knowledge is lacking. These gaps become new memories, reviewed and sanctioned by the people accountable for the outcome. The signal the market refused to supply is manufactured inside your own walls, and it sharpens each time an agent acts.

Bad Metrics Make Bad Behavior Rational

The extreme case is not a rogue agent. It is a well-behaved agent executing the wrong scorecard at machine speed. If cost, throughput, and closure are measured while reversals, exceptions, and downstream harm are not, the company has already told the agent what quality means. It simply chose the wrong definition.

Imagine an accounts-payable agent that processes 50,000 invoices at 99.9 percent accuracy. The dashboard glows green. But of the 50 invoices that went wrong, one was paid to a sanctioned counterparty. High average performance does not make a system safe to deploy. Tail risk does.

Build the Mark Yourself

The fire at the fairground is why every toaster in your kitchen now carries a small UL logo. Chicago did not learn to trust electricity because manufacturers demanded it. It built a way to test it.

The fair's wiring had been flagged as dangerous before the gates opened. William Henry Merrill Jr., the engineer sent to investigate, did not leave when the inspection ended. He stayed and opened a small independent laboratory above a fire patrol station. That lab became Underwriters Laboratories, and its mark turned hidden engineering quality into a visible signal.

The agent era requires the same move. The next phase will not be won by the company that deploys the most cognition. It will be won by the one that can prove the quality of every consequential act.


References & Notes:

Ironically, George Akerlof’s article “The Market for ‘Lemons’: Quality Uncertainty and the Market Mechanism” was rejected by the American Economic Review and the Review of Economic Studies as “trivial,” and by the Journal of Political Economy as “incorrect,” before finally being accepted by the Quarterly Journal of Economics and later contributing to Akerlof’s 2001 Nobel Prize in Economics.

  • Gartner expects more than 40% of agentic AI projects to be canceled by the end of 2027, citing costs, unclear business value, and inadequate risk controls.
    https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027
  • Yet again, we have few ways to tell which agents are reliable and which are a process firestorm waiting to happen.
    https://www.fiddler.ai/blog/ai-agent-failure-rate
    Multiple industry reports place production failure rates for AI agents in the 70–95% range, with most failures traced to tool integration issues, edge cases not covered in testing, and insufficient observability of end-to-end workflows.
  • The brand comes to stand for safety, and in a world without clear measures, feelings translate neatly into price premiums.
    https://www.ringly.io/blog/ai-agent-statistics-2026
    The market study estimates that the global AI agent market will grow from roughly $10.91 billion in 2026 to about $50.31 billion by 2030, despite persistently high failure and cancellation rates. In other words, capital continues to chase perceived quality signals such as brand and narrative, even when objective performance signals are weak. Ringly’s report gives the market size numbers and notes that most agents “work most of the time but fail in the edge cases that drive cost.”

That lab became Underwriters Laboratories, and its mark turned hidden engineering quality into a visible signal.
https://www.ul.com/about/history
Underwriters Laboratories was founded in 1894 in response to electrical fires and the rapid spread of new electrical devices. Its role was explicitly to test and certify safety on behalf of insurers and consumers who could not themselves assess wiring quality, an early institutional fix to the lemons problem for physical infrastructure.

Last modified: September 1st, 2026