Back to Research

TECH EVOLUTION // 05

Tech EvolutionAug 13, 202610 min

Smarter Models, Same Spreadsheet

Model capability rose fast and verifiably — Humanity's Last Exam from under 10% to 38.3% in a year, SWE-bench Verified to 76.8% by February 2026 — while value capture inside organizations did not keep pace; buyer data now demands proof of P&L over productivity anecdote, moving the race from who has the smartest model to who can show the result.

Two numbers came out of the same twelve months, and almost nobody puts them next to each other. One of them holds up. The other is the reason I now read the footnotes.

The first: on Humanity's Last Exam — a 2,700-question benchmark built specifically to be too hard for AI — frontier model accuracy went from under 10% to 38.3% between 2024 and 2025, a 30-percentage-point jump in a single year (Stanford HAI, 2026 AI Index Report, Technical Performance; data from lastexam.ai). On SWE-bench Verified — real GitHub issues, real codebases, a working patch or nothing — the top model reached roughly 76.8% by February 2026, with a tight cluster of competitors between 70% and 76% (SWE-bench Leaderboard, 2026), up from a field that was clustered around 60% the year before. That's not a plateau. That's a curve doing what curves do when nobody's found the ceiling yet.

The second number is the one everybody quoted: 95% of enterprises getting zero return on $30–40 billion in GenAI spending (MIT Project NANDA, "The GenAI Divide: State of AI in Business 2025" — deck dated July 2025, circulated publicly that August; 300+ initiatives reviewed, 52 structured interviews, 153 senior-leader survey responses). It ran in Forbes, Fortune and Inc. within days.

It also doesn't survive opening the report. The 5% figure appears in one section and refers specifically to custom enterprise AI tools that were "successfully implemented" — where success is defined as "causing a marked and sustained productivity and/or P&L impact." As Wharton's Kevin Werbach put it after reading the document several times, "'unsuccessful' explicitly does not mean 'zero returns'" (Futuriom, August 26, 2025). A narrow subset that didn't clear a high bar got reported as near-total failure of a much broader category, and the report never shows the arithmetic connecting them.

So the honest version of the gap is smaller than "95%" and much harder to headline. It also doesn't go away — and the most useful evidence for it turns out to be a number that neither side of that argument collected. That gap is the actual story of applied AI in 2026 — not "is it getting smarter" (yes, measurably, quickly), but "why doesn't smarter reliably turn into value," and what's starting to change about that.

Where we actually are

The capability side is the least controversial part of this. Stanford's AI Index scales benchmark results against human baseline performance and finds that on tasks like PhD-level science questions (GPQA Diamond), competition math, and multimodal reasoning, frontier models have reached or passed the human bar in the last two years. Coding and computer-use agents are earlier on the curve but moving fastest: the Index's own framing is that SWE-bench Verified performance rose "from approximately 60% in 2024 to close to 100% in 2025" — measured as a percentage of human baseline, not raw percent-solved. The raw leaderboard number for the best model in February 2026 was 76.8%. Both numbers are true and they're measuring different things; the "close to 100%" is relative to a human comparison point, and the 76.8% is the literal share of real-world GitHub issues a top agent resolves end to end. I'd rather cite the number that survives someone opening the source than the one that reads better in a headline, and the honest version is still remarkable: roughly a 17-point jump in raw resolution rate in about a year, on a benchmark that didn't exist to be easy.

Meanwhile the progress is jagged, not uniform, which matters for what comes next. ClockBench tested top models on reading analog clocks — 180 clock designs, 720 questions. Humans get it right 90.1% of the time. The best model in March 2026, GPT-5.4 High, got 50.6%, and when it was wrong, its errors ran one to three hours off; human errors ran about three minutes off (Stanford AI Index, citing Safar, 2025). A system that can pass expert-level science exams and still misread a clock face isn't broken — it's a reminder that "smart" in these systems is task-shaped, not general, and the tasks that map cleanly to structured, verifiable formats (code, math, closed-ended QA) are exactly the ones improving fastest. That's not a coincidence; it's the same reason SWE-bench and HLE make good benchmarks and bad proxies for "can this run my invoicing process."

Which is the other half of where we are — and it deserves the same scrutiny the benchmark numbers just got. NANDA stands for Networked Agents and Decentralized AI. It's a Media Lab spinoff building protocols for autonomous agent infrastructure, backed by companies with a direct stake in that architecture — several of whose employees are among the authors. A study concluding that today's centralized, off-the-shelf deployments mostly don't pay off is not a neutral observation coming from that address. So I'll take its qualitative finding and leave the headline: the failure mode it describes isn't model quality, it's the absence of a learning loop. Off-the-shelf tools plateau because they don't adapt to a business's actual workflow; the organizations that did capture value built or bought systems that remember context, integrate with existing process, and get corrected over time. That part matches what practitioners describe and what I've watched in my own builds.

Now set it next to a Microsoft-commissioned IDC survey from November 2024 — 4,000 respondents, average reported ROI of $3.70 per dollar spent, $10 for top performers (Microsoft/IDC, November 2024). The tempting move is to pick whichever flatters your priors. The accurate one is to notice that both came from parties holding a position: one by the vendor whose product is being evaluated, collecting self-reported perception; the other by a project promoting the architecture it argues the market is missing. Neither is worthless, and neither settles it. What moves the question forward is a different kind of number altogether — not one more verdict on whether AI works, but a measurement of what buyers changed their minds about.

The trajectory

That measurement arrived in February 2026, and it's worth more than either survey above because it's a delta, not a verdict — the same instrument, two waves, asking buyers what they now count as return. Futurum Group's 1H 2026 survey of 830 enterprise IT decision-makers found that "productivity gains" — the default justification for GenAI spend through 2024 and 2025 — fell from 23.8% to 18.0% as the top-cited ROI metric, while direct financial impact (revenue growth plus profitability, split out as its own category for the first time) nearly doubled to 21.7% of primary responses (Futurum Group, "Enterprise AI ROI Shifts as Agentic Priorities Surge," Feb 17, 2026). Their research director's framing: "the productivity argument was the right metric for the GenAI pilot phase, but the market has matured. Enterprises are now demanding that every AI capability connect directly to revenue growth or margin improvement." Read plainly, that's buyers admitting the productivity-anecdote era didn't hold up to a finance department's questions, and moving the goalposts to something harder to fake.

The same survey found Autonomous Agents/Agentic AI surged 31.5% year-over-year as the top technology priority (17.1% of respondents, up from 13.0% six months earlier). Extrapolated in a straight line, that means the next eighteen months of "applied AI evolution" is mostly going to be an accountability story, not a capability story: capability will keep compounding on structured tasks roughly as fast as it has, while the market's actual demand shifts to "prove the P&L line," and the vendors and builders who can show that proof — not the ones with the highest benchmark score — win the next round of budget. The limit on this extrapolation is real: a single mid-cycle survey from one research firm is a signal, not a law, and "enterprises say they'll demand harder ROI proof" is a stated intention, which is cheaper than a measured outcome. Whether 2027's version of the NANDA study finds the divide narrowed or just relabeled is an open question, not a settled one — and it connects directly to the gap I wrote about in Code Got Cheap, Trust Didn't: every time generation capability outruns the verification layer around it, the bottleneck migrates somewhere unmeasured — and unmeasured territory is exactly where a badly-sourced statistic can travel for a year before anyone opens the PDF.

What if... (labeled speculation — these are scenarios, not forecasts)

What if the 5% club's pattern gets productized? NANDA's own finding is that the winners built narrow, learning, workflow-specific systems rather than deploying general chat tools. If that pattern gets packaged into reusable software instead of staying bespoke-consulting-hours, the divide could narrow fast in 2027 — not because models got smarter, but because "integration" stopped being custom labor. What would need to be true: someone ships an integration layer that's genuinely reusable across companies, not just a services engagement wearing a product's logo.

What if the benchmark curve and the deployment curve stay permanently decoupled? The tasks compounding fastest (code, closed-form math, structured QA) may simply not be the tasks that make up most of real operational work, and the ClockBench-style long tail of "trivial for humans, unreliable for models" may persist for years even as headline benchmarks approach saturation. In this world, "AI capability" keeps looking exponential in the reports that measure what's easy to measure, while deployable reliability grows much closer to linear. What would need to be true: benchmark design keeps optimizing for what's gradeable, and the gap between gradeable and valuable keeps being underpriced by everyone writing headlines about it — including, some weeks, me.

What if the divide is partly a measurement artifact of timing, not substance? NANDA's data was collected January–June 2025, near the peak of the pilot-first wave. A 2027 cohort that adopts narrow, custom-built systems from day one — because the 2025 failure mode is now well-documented — might simply never generate "failed pilot" data in the first place, closing the divide by changing who gets counted rather than by anyone learning a lesson. That's a real possibility, and it's a second reason — on top of the sourcing problem above — not to treat that figure as a fixed property of enterprise AI rather than a snapshot of one immature moment, measured loosely.

What I'm doing about it

I don't run a general-purpose AI deployment — I run several narrow ones, and the NANDA finding matches what I've watched happen in my own build. The most productized asset in my portfolio, agent_factory, is not a chatbot wrapped around a business problem; it's a fixed pipeline — Scout finds signal, Producer drafts, a Curator step gates output against a fixed bar before anything ships, Publisher executes — running at zero marginal token cost on a fixed compute plan. It's boring by design, narrow by design, and it has a review gate that nothing skips. That's the same shape NANDA describes in its 5%: not "smarter model," but a system built for one process, with memory of what worked last time and a checkpoint before anything goes live.

The article you're reading came out of the same discipline. It doesn't publish on a schedule alone — it publishes after a separate reviewer checks the sourcing against the primary document, not the summary, and sends it back if a number doesn't survive that check. That loop is slower than "generate and post." It's also the reason the second number in this article got rewritten before you read it rather than after — the draft that reached review led with "95% of organizations report zero measurable P&L return," and the review is where that sentence died. The capability curve is real and I use it every week. The value only shows up when something narrower than "AI" — a scoped process with a gate on it — sits between the model and the outcome.

Next week: back to whichever field has the sharpest new signal when Tuesday's research starts — cybersecurity or software development are both due for a fresh angle. Research home: blackicelabs.ca

Written by Levon Azevedo

Get the next one by email.

Subscribe