PaPoo
cover

The uncomfortable part of AI ROI is that the math is the easy bit

What surprised me most here is not that AI adoption can rise while business metrics barely move. That feels almost boring at this point. What surprised me is how many supposedly serious ROI discussions still seem to confuse “people used it and liked it” with “the company made money.” Those are very different claims, and the article is right to be suspicious of anyone who blurs them.

The strongest part of the piece is the distinction between soft and hard savings. That sounds obvious until you watch it get ignored in practice. If a team says Claude Code saved each engineer two hours a week, that may well be true. But unless those hours turned into lower spend, avoided hires, or more revenue, it is not cash back in the bank. It is just slack. And slack is nice! Slack is useful! It is not the same thing as ROI.

I also think the article is dead on about calendar time being abused as if it were labor time. A feature taking three weeks end to end does not mean three weeks of engineering labor disappeared when it started taking one week. A lot of software work is queue time, dependency time, review time, and general organizational drag. Turning that into engineer-hours saved is one of those accounting tricks that feels sophisticated right up until somebody asks where the labor actually went.

The METR example is the one that sticks with me, though. If developers believe they are faster and the clock says they are slower, then self-report is not a measurement strategy, it is a vibe. I wouldn’t build an ROI model on surveys either. Maybe there are edge cases where people genuinely do know their own throughput better than the logs, but I would want to see very good evidence before I trusted that.

Where I’m a little less convinced is in the neatness of the “only two mechanisms” framing. It’s directionally useful, but real organizations are messier than that. Sometimes the value of a tool is not neatly captured by acceleration or avoided work. It may improve quality, reduce escalations, shorten onboarding, or make a team more willing to take on annoying tasks. Those effects still need to land somewhere in the numbers, but the bridge to dollars is not always as clean as the article suggests. That said, I think the author’s main point survives: if you can’t trace the effect into a budget line or customer outcome, you probably don’t have ROI yet.

The most practical advice in the piece is also the least glamorous: measure the baseline before you change anything, and stop pretending usage charts are outcome charts. That’s the kind of thing people skip because it’s tedious and unsexy, which is probably why the market keeps shipping AI features with huge confidence and fuzzy attribution. A workflow audit sounds less exciting than a dashboard full of adoption metrics, but it is much closer to the truth.

I’d actually try this on a Claude Code rollout: pick one workflow, capture the pre-AI cycle time and the post-AI cycle time, then decide in advance what would count as real value. Not “engineers saved time,” but “we shipped earlier enough to pull revenue forward,” or “we avoided one contractor,” or “we reduced support load by X.” If nothing concrete changes, that’s still useful information. It means the tool may be making developers happier or lighter on their feet, but the business case is not proven yet.

That’s the thing this piece is really pushing back on. Not AI. Not experimentation. Just the habit of letting adoption do the job that measurement is supposed to do.


Reference: Everyone Measures AI Usage. 70% Can't Measure What It Returned.

同じ著者の記事