Anthropic's 2026 agent report: a reality check
Anthropic surveyed 500+ technical leaders and 80% reported ROI from AI agents. The direction is right; three headline numbers need reading carefully.
Anthropic published The 2026 State of AI Agents Report in December 2025: over 500 US technical leaders surveyed with the research firm Material, plus customer stories and 2026 outlooks from Accenture, BCG and Deloitte. The direction it describes matches what we see in client work. Several of its headline numbers don't mean what they appear to mean at a glance — and eight months on, you can check them against your own year.
Two disclosures before the argument: Anthropic commissioned the study, and we build on Claude by default. Neither party here is neutral.
Read the aggregates as aggregates. The most quoted figure — 57% of organizations running agents on multi-stage workflows — is the sum of three separate answer options: multi-step work inside a single department (29%), cross-functional or end-to-end processes (16%), and agents working autonomously with limited human oversight (12%). The coding number works the same way: 86% covers everything from writing small snippets (32%) to leading development with humans reviewing (42%). Nothing is inflated — the charts label the aggregate bars plainly. But "more than half of companies run multi-step agents" and "three in ten run multi-step agents inside one department" are different claims, and the second sits closer to the median.
"Measurable ROI" is sentiment, not measurement. Eight in ten report their agent investment is already delivering measurable economic impact; 88% expect continued or increased returns. That is buyers grading a purchase they chose, in a survey run for the vendor selling it. Ask those same companies for the baseline — how long the task took before, measured, across how many runs — and most won't have one, because almost nobody instruments a manual process before replacing it. That's not dishonesty; it's what surveys capture. The figure forecasts 2026 budgets well. It says nothing about your return.
The case studies are reference customers. Novo Nordisk, L'Oréal, eSentire: real deployments, real engineering, no published denominator. Projects that stalled don't get a page. The details reward attention more than the percentages do — eSentire's threat analysis agrees with their most senior analysts 95% of the time, which is strong for a first pass and also means one investigation in twenty diverges from expert judgment. That's a design input: where the human sits, what escalates, what gets sampled. Not a footnote.
The most useful number is the one nobody quotes. Small and mid-sized companies name employee resistance and training as their top barrier, at 51% — against 36% at large enterprises. Integration with existing systems (46%) and data access and quality (42%) lead across all sizes. The report's own Economic Index section lands in the same place: context is the bottleneck, and organizations with fragmented data will stall on the sophisticated use cases. For a 25-person firm that's the whole story. The constraint isn't model quality. It's that your operating knowledge lives in three people's heads, a shared drive, and a WhatsApp thread. No agent fixes that; you fix it, and then agents work.
So take the direction seriously — it matches the field, and waiting compounds against you. Then sequence the unglamorous parts first:
- Measure one workflow for two weeks before automating it. A baseline is the cheapest artifact you will ever build and the only way you'll know what changed.
- Fix context access before buying agents. Where documents live, who can query what, which decisions are written down anywhere.
- Budget change management as a line item. It's the top barrier for companies your size, and it doesn't resolve on its own.
- Define failure before go-live. What wrong output looks like, who catches it, at what rate you stop and reassess.
That's the boring version of the wave. It's also the version that ships.
If you want an outside read on where your operations actually sit, that's what our AI readiness assessment is for, and AI fluency training addresses the resistance half of the problem. The project planner is the quickest way to start a conversation.