JS
Back to All Articles

Four Reports Studied the Same Year and Disagreed About Everything

July 28, 2026  ·  15 min read

Four Reports Studied the Same Year and Disagreed About Everything

Four serious research organizations spent the same twelve months studying what happens when companies deploy AI. They came back with answers that cannot all be true at once. If you are approving an AI budget this quarter, every number you have been handed is wrong in a different direction. The disagreement is the most useful thing in the pile.

Here is what I read.

Stanford's Digital Economy Lab published a study of 51 AI deployments across 41 organizations, all live, all in production, all producing measured value. Flexential surveyed 354 IT decision-makers at companies above $100M in revenue. A team from Cambridge, MIT, Harvard and Stanford audited 30 deployed AI agent products across 45 disclosure fields each. Anthropic surveyed more than 500 technical leaders about where agent adoption actually stands.

Two of those organizations have nothing to sell you. Two of them do. Hold that thought, because it turns out to be the most reliable filter in the stack.

They disagree about whether any of this works at all

Anthropic reports that 80% of leaders say their agent investments already deliver measurable financial impact.

Flexential reports roughly 20% are seeing returns today.

MIT's NANDA initiative reports that 95% of generative AI pilots produce no measurable financial impact at all.

Stanford reports 100%, because they only studied deployments that worked. They say so openly, in the methodology section, in the limitations section, and again in the conclusion.

Somewhere between 5% and 80% of AI projects are paying off, depending entirely on who is asking and what they sell.

These numbers are not in conflict. The denominators are, and nobody publishes theirs. MIT counts pilots. Anthropic counts organizations. A company can run twenty failed pilots and one success, and answer honestly that yes, they see financial impact. MIT scores that company a 95% failure. Anthropic scores it a win. Both are telling the truth about different things.

So when a vendor deck lands on your desk with a number on it, the question is not whether the number is real. The question is what sits underneath the division sign. Ask it out loud in the meeting. Watch how long it takes to get an answer.

They disagree about what is actually stopping you

Stanford found that 77% of the hardest challenges practitioners faced were invisible: change management, data quality, process redesign. Not technology. Executive sponsorship was the single most-cited accelerator at 43%. Their conclusion is blunt about it. The technology works. The challenge is everything else.

Flexential found that at the top of the market, executive support is a barrier for 5% of respondents and budget for 1%. Infrastructure leads at 40%, security at 33%, skilled staff at 22%.

Read fast, those contradict. Read properly, they are a relay race.

Flexential's respondents already won the organizational fight, which is why physics is what remains. Power, chips, latency, fiber. Stanford's sample is still in the middle of the organizational fight, which is why sponsorship dominates.

There is also a measurement artifact worth naming, because it will save you from misreading a survey again. Flexential asked what blocks you. Stanford asked what accelerated you. Having a sponsor is not a barrier. Lacking one is. Same variable, opposite questions, incompatible-looking answers.

For a company between 50 and 500 people, this is the useful read: you get to fight the organizational battle without ever hitting the physics wall, because you will rent your compute instead of building it. The problems slowing down the largest companies in the world are not going to be your problems.

They disagree about how much rope to give the machine, and this one is a real fight

Stanford classified every deployment by oversight model. Escalation means AI handles 80% or more on its own and humans review only the exceptions. Approval means a human signs off on every output before anything happens.

Escalation delivered 71% median productivity gains. Approval delivered 30%.

That is the loudest number in the entire study, and the obvious reading is: stop gating everything, the money is in letting go.

Now hold that next to the Agent Index, which audited what the agent products on the market actually disclose. Enterprise agent platforms have a design and deployment split. A human configures the agent under close supervision on a visual canvas, and then the deployed agent runs on event triggers with no human in the loop at execution time, permanently. Guardrails ship as optional modules. Of 30 audited products, 21 have no documented default behavior for telling a person they are talking to an AI. Four have no fine-grained stop controls, meaning the only way to halt the thing is to retract the entire deployment. Four have agent-specific safety evaluations. The paper names the pattern directly: platforms advertise their compliance certifications while agent-specific guardrails become the customer's responsibility.

So Stanford says loosen the reins because that is where the returns live. The Agent Index says the reins came off already, you were not consulted, and nobody built the brakes.

You have already made the autonomy decision Stanford is telling you to make. You made it by accident, and you made it without the exception handling.

Two days ago I tried to cancel a DocuSign account

I opened the chat window. I explained what I needed.

The bot told me to contact support.

I typed back: you are support.

It did not have anywhere to go with that. There was no path in the system for the thing I was asking, and no path for escalating to a human who had one. So I closed the window and figured out the cancellation myself, clicking through settings until I found it.

DocuSign is a company whose entire business is digital workflow infrastructure. Signature routing, approval chains, conditional logic, audit trails. They sell the machinery of getting a process from start to finish without a person having to hold it together. They employ excellent engineers.

And their customer-facing agent could not handle a customer asking to leave.

That is not a governance failure. Nobody at DocuSign forgot to write a policy. That is a build failure, and it is the most common one I see.

The gap is engineering, and it is a specific kind of engineering

Here is the pattern. A company decides it needs an agent. It has engineers. Engineers build things. So the engineers build it, because they can, not because they have ever built one before.

The reports back this up from four directions.

Anthropic's report includes a case from L'Oréal that is the cleanest evidence in the stack. They built conversational analytics two ways. A single-model approach reached 90% accuracy. An architecture that orchestrated specialized agents reached 99.9%. Same problem. Same underlying models. One hundred times fewer errors. The entire difference was how it was designed. Treat the specific numbers as vendor-published, since Anthropic sells the model and the customer is theirs. The direction of the finding still holds, and it is the whole argument in one line: architecture, not model choice.

Stanford's failure appendix is more direct than I expected. One of the six root causes is a talent and sponsorship gap, at 12% of cases. The remediation column says to create dedicated roles and not simply retrain existing staff. A Stanford research team is telling you that pointing your current engineers at an agent project is a documented way to fail.

Another root cause, at 16%, is technology that broke or was not mature enough. Look at how companies fixed it: modular frameworks, hybrid designs where AI carries 80% and humans refine the last 20%, and dual-model validation before trusting a single output. Every one of those is a practice, not a product. The technology was not too immature. Nobody applied the practices.

And the Agent Index gives you the inversion that should end the internal debate. Four of thirty commercial agent products publish agent-specific safety evaluations. These are companies whose entire business is building agents. If the specialists are not running agent-level evals, the odds your internal team wrote any are close to zero.

Flexential prices the scarcity from the other side. The skilled staff barrier doubled from 10% to 22%, and 76% of organizations now pay a salary premium of 10% or more for AI skills. Markets do not pay premiums for things that are easy to find.

A telecom executive quoted in the Stanford study named the missing role better than I could. Their number one issue was not a shortage of AI people or a shortage of process people. It was a shortage of people who understand the process and the AI and can put the two together.

No agent is ever finished, and that is the part nobody budgets for

I update my own system almost every day.

I run an AI operating system for my businesses. Agents that handle scheduling, drafting, meeting intelligence, follow-up, financial reconciliation. It has been live for months and it is genuinely good. It is also never done. Last week I fixed a failure class where a dropped connection meant a real person went a week without hearing from me. That was not a catastrophe. That was Tuesday.

This is the piece non-technical leaders consistently miss, and it is the difference between the two kinds of bad first attempt.

Stanford found that 61% of their successful deployments had a failed AI project before the one that worked. Every single case where they could identify the methodology used an iterative approach. Not one used waterfall. Failure first is not the exception in this data. It is the path.

So a rough first version is not evidence of incompetence. Shipping a rough version and then walking away is. That is the distinction, and it is the whole thing.

Agent systems are software. You QA them. You test them. You find the case nobody anticipated, and then you find the next one. This is not a phase you get through. It is the permanent operating condition, exactly as it is for every other piece of software your company runs, except that the failure modes are stranger and the system will fail confidently instead of throwing an error.

Which means someone has to own it. Not a project sponsor. An operator. Somebody whose actual job is monitoring the agent systems, running audits on what they produced, debugging the failures, updating the scaffolding around the model as the work changes, and periodically going back through the accumulated rules and compacting them so the thing does not collapse under its own instructions.

An agent is not a project with an end date. It is a system with an operator. If you cannot name the person, you do not have an agent. You have an unattended process talking to your customers.

Nobody puts that line in the business case. It is the line that determines whether the business case is real.

Scale will not save you, and it might be the problem

The obvious objection to everything above is budget. Surely a company with a real engineering organization and real money does this properly.

DocuSign is a multi-billion-dollar company. That is the point of the story.

Meanwhile, look at which deployments Stanford labeled mid-market in their own sample. A regional supermarket chain with about two dozen stores, running at roughly half the industry benchmark margin with almost no negotiating power against suppliers, put an autonomous agent on procurement and cut waste 40%, cut stockouts 80%, and doubled EBITDA margin. A call center company embedded agents into the product itself and won more than twenty new projects, ending up benchmarked among the top four for AI in customer experience against companies that were built AI-native from day one. A six-person security team went from processing 1,500 alerts a month to 40,000, redeployed four and a half people to threat hunting, and laid off nobody.

Those are not scrappy underdog stories. They are companies that made a good architecture decision and then kept operating the thing.

If budget and headcount were the constraint, DocuSign would have the better chatbot. They do not. Which means the constraint is judgment, and judgment is purchasable at any size.

Here is the uncomfortable version for the larger organization. A company of 200 people knows it has to go find someone who has built one of these before. A company of 10,000 has enough engineers to build it badly in-house, and enough internal politics to keep defending it after it ships. The large engineering organization is not always an advantage here. Sometimes it is the reason nobody ever said the words "we have not done this before."

What to do with this before Friday

Find your agents and write down what each one is allowed to do without a human. Not what the vendor demo showed. What is deployed. For each one, answer three questions: who can stop it mid-execution, what happens when it is wrong, and does it tell the person on the other end that it is not a person. If you cannot answer all three for any agent on the list, that is not a documentation gap. That is the finding.

Name the operator. One human being, by name, whose job includes monitoring and auditing those systems on an ongoing basis. Not the project sponsor. Not the vendor. If the name is blank, you have deployed something and abandoned it, and the deployment date was the last day anyone looked at it.

Ask your agent vendor for their agent-specific safety evaluations. Model-level evaluations do not count, and they will try to give you those. You want the evaluations of the agent, the thing that takes actions in your business. Based on the Agent Index numbers, there is roughly an 87% chance the answer is that they do not have any. That answer is not a reason to panic. It is a reason to understand that the guardrails are yours to build, and to budget accordingly.

Four reports read the same year and disagreed about nearly everything. They agree, quietly and without stating it, on one thing: the hard part was never the model.

Somewhere in your organization, a system you approved is talking to a customer right now. The only question worth asking is whether anyone is still listening to it.

Ready to scale with clarity?

I take on 3–4 new clients per quarter.