Field Notes · Production AI

Three rooms, one question: what August taught me about why AI does not pay for itself

Three sessions, three audiences, ten days. All three asked the same question without knowing it, and the answer explains why AI does not pay off.

Anil Prasad · August 25, 2026 · 9 min read
Title card reading Three rooms, one question, from Field Notes: Production AI by Anil Prasad, Global Product Leadership Festival 2026.

In ten days this month I spoke three times at the Global Product Leadership Festival. An AMA on portfolios for AI product work. A fireside chat on leading AI product teams. A masterclass on monetizing AI products.

Three formats. Three audiences. Three questions that on paper have nothing to do with each other.

They turned out to be the same question.

What each room thought it was asking

The portfolio room wanted to know what to put in a case study. Underneath that: how do I show a reviewer that I am good at this?

The leadership room wanted to know how product, engineering, data and design work together on AI. Underneath that: why do our teams keep disagreeing at the worst possible moment?

The monetization room wanted to know what to charge. Underneath that: how do we make this pay?

Here is what all three have in common, and I did not see it clearly until the third session.

In every case, the thing that was missing was a defensible record of what happened and why.

The candidate could describe what they shipped and not how they decided. The team could describe four different definitions of "done" and had never written any of them down. The vendor could describe a resolution and could not define it in terms a buyer would accept.

Same failure. Three costumes.

Three sessions compared: an AMA on portfolios, a fireside chat on team leadership, and a masterclass on monetization. In each, the stated question differed but the underlying failure was the same missing record of what happened and why.

What I got wrong in the first session

I went into the AMA planning to talk about portfolios and spent the first ten minutes apologising for not being a designer. That was the wrong instinct. The useful thing I had was not design craft, it was that I sit on the other side of the table and review this work.

I have reviewed portfolios where every project succeeded. They are the least useful kind. A reviewer already assumes you shipped things. What they cannot tell from a gallery of wins is how you decide, which is the only thing a forty-five minute conversation is actually trying to establish.

The framing I landed on afterwards, which I would have opened with if I had thought of it sooner, is four questions.

Proof of problem. What was true before, with a number attached. Not "users were frustrated" but "resolution took eleven days and 40% of tickets were reopened."

Proof of decision. What you chose not to build, and why. Anyone can describe what they shipped. Very few can describe the option they killed.

Proof of behaviour. What the system actually did in production, including when it was wrong. If you cannot describe a case where your AI feature produced a bad output and what you did about it, an interviewer assumes you were not close enough to it.

Proof of consequence. What moved, and what you would hand somebody who is allowed to check.

Only the third one is hard to fake, which is why almost nobody includes it.

The Four Proofs framework: proof of problem, proof of decision, proof of behaviour, proof of consequence.

Does the evidence problem actually cost money?

It is easy to file this under governance and move on. That would be a mistake, and the numbers say so.

McKinsey surveyed 1,993 organisations across 105 nations, fielded between 25 June and 29 July last year. 88% report regular AI use in at least one business function. 39% can attribute any EBIT impact to AI at all. 6% see AI drive five percent or more of enterprise EBIT.

Read that middle number carefully, because it is usually misread. It does not say 39% are failing. It says 61% cannot point to a single point of profit from something they have already deployed.

Every company in the 88 has working models. Model quality is not what separates them from the 6.

88 percent of organisations report AI use, 39 percent can attribute any EBIT impact, and 6 percent see AI drive 5 percent or more of enterprise EBIT. McKinsey, n equals 1,993 across 105 nations.

Gartner reached the same conclusion from a different direction. In June last year they projected that over 40% of agentic AI projects will be cancelled by the end of 2027, and named three causes: escalating costs, unclear business value, inadequate risk controls. Not one is a model problem. All three are questions about whether you can show your work.

What this looks like at the pricing layer

The monetization session made the connection concrete in a way the other two could not.

Look at what the AI customer service market currently charges. One vendor bills $0.99 per resolution and nothing if the conversation escalates. Another bills $2.00 per conversation whether or not it was resolved. Same category, same buyer, roughly the same job.

That is not a pricing difference. One sells a result. The other sells capacity and lets the customer carry the failure rate.

Now the part that should interest anyone who has ever signed a software contract. A leading vendor reports an average resolution rate of 76% across more than eight thousand customers. Independent reports place the same metric between 42 and 50%. On a per-resolution price, that spread is the difference between a good deal and a bad one, and it exists because "resolution" was never defined in a way both parties could verify.

A competitor's response was instructive. Rather than cutting price, Zendesk restructured in May to bill only on a verified resolution: one confirmed by a separate evaluation model within 72 hours. Everything else is free.

They did not compete on price. They removed the argument by adding an auditor.

That is a product decision, made by a product team, showing up directly in the revenue line. Provenance stopped being a compliance topic somewhere around that release and became a pricing feature.

Why outcome pricing stays rare for longer than people expect

The published expectation is that outcome-based pricing goes from roughly 5% of companies today to about 25% by 2028. I would take the under, and here is my reasoning rather than my conclusion.

Outcome pricing transfers model risk from the buyer to the vendor. To carry that risk you need three things at once: an outcome you can measure cheaply, a success rate stable enough to forecast, and an evidence trail good enough to defend a disputed charge nine months later.

Most companies have none of the three. Some have the first. Very few have the third, and the third is the one nobody is building because it does not show up in a demo.

So my expectation is that hybrid pricing keeps winning, not because it is better but because it is the only model you can run without the evidence layer. Hybrid went from 27% to 41% of companies in a single year. That is not enthusiasm for hybrid. It is the market routing around a capability gap.

I could be wrong about this. If someone builds a genuinely cheap, portable outcome-verification layer, the 25% number becomes conservative rather than optimistic. I do not see anyone doing it yet, but I would not have predicted the Zendesk verified-resolution move either.

Where I learned this the expensive way

I have been on the wrong side of this.

Running AI platforms for healthcare revenue cycle, one of our agents silently approved $2.4 million in non-covered procedures over roughly six weeks. The number is not the lesson.

A CFO found it in a month-end review. Our audit pipeline never did.

Monitoring said the system was up but could not say what it had done. 2.4 million dollars in non-covered procedures approved over six weeks, caught by a CFO in a month-end review rather than by the audit pipeline.

Uptime, latency, throughput, error rates: green the entire time. We had monitoring. We did not have evidence. Those are different things and almost every team I meet has confused them.

The connection I missed for years is the one this month made obvious. The capability that would have caught that in week one is the same capability you need to bill on outcomes, defend a disputed invoice, or answer a regulator. One investment. Most organisations have it filed under three separate budgets and fund none of them properly.

One correction worth making

Since we are on the subject of checking claims, two that circulate constantly and should not.

The "95% of AI pilots fail" figure comes from a report whose zero-return finding rests on 52 interviews, and which describes that finding itself as directionally accurate based on individual interviews rather than official company reporting.

And the belief that AI has destroyed software margins does not survive contact with the data. Aleph and Benchmarkit published FY2025 actuals across 342 companies on 1 June. Median software gross margin: 80%, and it held between 79 and 81% for four straight years. What is true is narrower and more useful: usage-based pricing models run at a 62% median and AI-native companies around 52%. Margin is now a consequence of the pricing model you choose rather than of the sector you happen to be in.

Median software gross margin 80 percent, total revenue 76 percent, usage-based pricing models 62 percent, AI-native companies 52 percent. Aleph and Benchmarkit, n equals 342 companies, FY2025 actuals.

The widely quoted "AI margins are 50 to 60%" figure, incidentally, originates in a 2020 essay. It gets recycled as current analysis six years later.

What I would have said differently

Two things.

In the fireside chat I said that most cross-functional AI failures come down to four functions holding four definitions of "done." That is true and it is incomplete. The part I left out is that the four definitions are all correct inside their own frame. Product is right that done means shipped. Engineering is right that done means stable under load. Nobody is being unreasonable, which is exactly why the argument never happens until launch.

The second thing is about the layoff data I quoted. US tech has shed roughly 28% of product management headcount from the 2022 peak, and the cuts are heavily skewed by level: VP-level roles down around 38%, senior individual contributors down around 21% and the most resilient group in the set. I framed that as encouraging for practitioners. Having sat with it for a few days I think the honest framing is narrower: the layer being compressed is coordination, and if your value is routing information between other people, that is a real warning rather than a reassurance. I should have said that plainly. These figures come from an industry aggregation rather than a controlled study, so treat the direction as more reliable than the decimals.

What I would do differently on Monday

Three things, none of which needs budget or permission.

Write down what would make you kill the AI feature, with a number and a date, before it launches. Teams that cannot write that sentence are not aligned. They are just not arguing yet.

Make one person accountable for the evaluation harness as a product, with a roadmap and a name attached. In every organisation I have worked in, evaluation is the actual bottleneck and it is the one thing no job description mentions.

Then measure time to reconstruct. If a customer disputes an output today, how long until you can show exactly how it was produced? It is measurable, nobody tracks it, and it predicts your next bad quarter better than accuracy does.

The question I did not prepare for

One attendee asked something I had not prepared for: what do you do when leadership does not actually want the evidence, because producing it would force a decision they have been avoiding?

I said the honest thing, which is that it stops being a tooling problem at that point and becomes a political one. Then I gave the only advice I have that works: build it small enough that nobody has to approve it, and wait for the quarter when someone senior needs an answer fast.

The question I am still sitting with

Across all three sessions the pattern held, but I am not certain about the cause. It may be that evidence infrastructure is genuinely hard. It may simply be that nobody owns it, so it never gets funded.

If you have built this properly, I would like to know which it was: did you solve it with tooling, or by making one person unreasonably accountable? I have watched the second work more often than I expected, and I am not sure that is a good thing.

What I am actually going to do about it

I run a quarterly series of anonymised AI incident post-mortems, because the $2.4M case taught me that the useful artifact is not the war story, it is the structure underneath it: what was instrumented, what was not, who found it, and how long reconstruction took.

If you have an incident you can talk about with the names removed, I would like to hear from you. The first issue is being assembled now. What I want most are the boring ones. Everybody publishes the dramatic failures. Almost nobody publishes the six-week silent drift, and that is the category that actually costs money.

Anil Prasad has spent 28 years building and running production data and AI systems in energy, healthcare revenue cycle, genomics and financial services. He writes Field Notes: Production AI on what breaks after an AI product has real customers.

Every figure here carries a source and a sample size. If one is wrong, tell me and the correction runs in the next edition.

Agent Governance LLM Observability Regulated AI Production GenAI Healthcare RCM AI Monetization Product Leadership
Have an incident you'd put on the record?
I run a quarterly anonymized post-mortem series on regulated AI failures — the boring six-week silent-drift kind, not the dramatic ones. Names removed. Submissions open now.
Submit a case → More essays