Klarna announced in early 2024 that its AI chatbot was doing the work of 700 people. They'd cut support staff and save $40 million a year. Then, somewhere between late 2025 and early 2026, they quietly started rehiring.

It wasn't the number itself that was wrong. The problem was that it answered a question the number couldn't actually answer. When you try to respond to "what's our AI ROI?" with a single multiplier across the whole company, this pattern repeats itself.

TL;DR
Vendor PR numbers Klarna's reversal Accuracy -19 percentage points outside task boundaries Workflow redesign reaches only 21% Task-level validation benchmarks

Why 700 jobs flipped in six months

Look back at Klarna's actual numbers and the answer's right there. The AI handled straightforward inquiries fine, but it kept hitting walls with complex payment disputes, fraud reports, and policy exceptions. Customers called back for the same issue multiple times. The chatbot struggled to calm upset customers. The company ended up going hybrid: AI handled the first response, humans dealt with exceptions and customer retention.

The real issue wasn't the "700 jobs" figure itself. It was that the figure never said which tasks produced it. On simple queries, maybe it really did equal 700 people's worth of work. The moment you scale that number across all customer service, though—where 20-30% of tasks need exceptions—the costs start leaking right back out.

Why '2.5x ROI' figures are dangerous

This isn't a Klarna-specific story. That's what makes it scary. Harvard Business School and BCG ran a field experiment with 758 consultants. On tasks within AI's strength zone (inside the boundary), AI delivered 25% faster results and 40% higher quality. Step slightly outside that boundary, and consultants using AI were 19 percentage points more likely to produce wrong answers. The team called this the "jagged frontier"—same workflow, some tasks where AI crushes it, the next task over where it's actually worse than humans.

The instant you flatten "our company's AI ROI is 2.5x" into one number, you're averaging together what the tool does well and what it doesn't. McKinsey found that 88% of organizations using generative AI say they're using it regularly. Only 21% actually redesigned workflows to make it work. That redesign gap? It's the biggest predictor of EBIT impact. MIT's research went harder: 95% of the generative AI pilots they studied didn't produce measurable profit improvement at all. BCG surveyed over 1,800 executives. 75% called AI a priority. 25% said they'd actually realized meaningful value.

Company-wide flattened numberTask-level validation benchmark
Typical example"AI cuts 700 jobs"Customer support, standardized inquiries -15% handling time
SourceCompany press releasePeer-reviewed field experiment / controlled RCT
RepeatabilityBreaks when exceptions mix inClear task boundaries make reproduction possible
Klarna's outcomeReversed with rehiring six months laterDeveloper productivity +55.8% holds up

So what can you actually tell your CFO?

Credible numbers exist. They just come from places with clear task boundaries. GitHub ran a controlled experiment with 95 pro developers split randomly. The Copilot group finished the same HTTP server build task 55.8% faster (1h 11m vs 2h 41m), with higher quality too (78% vs 70%). Statistically significant (p=0.0017).

These numbers share one trait: source credibility comes in tiers. A vendor's customer story (Klarna's 700 jobs) shouldn't sit in the same category as a randomized controlled trial (Copilot's 55.8%). Stack them by credibility and you get this.

CredibilitySource typeExample
① HighestPeer-reviewed field experimentHBS/BCG consultant experiment, GitHub Copilot RCT
② HighInvestor disclosureFinancial figures in earnings reports
③ MediumInternal operational caseSavings verified on internal dashboards
④ Medium-lowVendor customer story"Our customer saved 700 jobs"
⑤ LowSingle-executive surveyExecutive gut-feel response (use only cross-checked with other sources)

Klarna's number lands in tier ④—vendor customer story. That tier's for reference, not for board reports as-is. Copilot's 55.8% or the consultant experiment's 25%? That's tier ①, safe to cite much more directly.

Five steps to a defensible ROI matrix

  1. Break work into tasks
    Don't lump "customer service AI rollout" into one bucket. Separate "standardized inquiries" from "payment disputes" and measure each.
  2. Lock in baseline per task
    Record volume, time, cost, quality, and exception rates before deploying AI. Without this, you have no way to prove later impact.
  3. Tag source credibility
    Use the five-tier table above. Mark tier ④ and ⑤ numbers as "reference only" and keep them separate in reports.
  4. Don't mix forecast, annualized, and realized savings
    "$40M projected" and "cost actually cut this quarter" go on different lines. Mix them and Klarna's story repeats. It's also the CFO's job to stop pilots from dragging on due to sunk-cost bias.
  5. Note workflow redesign separately
    Same tool, different outcomes between teams that redesigned process vs. teams that just plugged it into old workflows. Track redesign as its own column.

Watch out

When you expand AI beyond where it works well, be especially careful. Using success in one task to justify pushing into exception-handling and emotional work can drop accuracy below human performance.

Common questions

We've got a pilot already running. How do we validate its ROI after the fact?

Without a pre-deployment baseline, perfect validation's impossible. Instead, find a similar task in a team or location that hasn't adopted AI yet and run a quasi-experimental comparison. If that's not possible, at least set a fresh task-level baseline starting now and compare quarter-over-quarter going forward.

Every department measures differently. How do we get them comparable?

Don't force one global metric. Lock in just five items as department-wide standards: volume, time, cost, quality, exception rate. Keep everything else task-specific. When those five match, departments' multipliers may differ, but credibility tiers stay comparable.

Can we just report vendor-supplied numbers straight to the board?

If you tag them properly, yes. Call it "vendor-cited tier ④" and note it's not reproduced internally. That way, if it flips like Klarna did, your credibility doesn't crater.

Is there a minimum ROI unit we can measure without redesigning workflows?

Yes. Tasks with naturally clear boundaries—code review, standardized support responses—show repeatable results even plugged into old workflows. Tasks with mixed exception-handling don't hold up without redesign.

How do we spot failure signals early in a pilot?

Watch repeat-inquiry rates and handoff-to-human rates. Same question coming back multiple times, or AI not reducing the rate it hands off to staff? That task isn't inside AI's boundary yet.

Go deeper if you want

Navigating the Jagged Technological Frontier Original research with 758 consultants showing which tasks AI gets right and which it doesn't aiinstitute.hbs.edu

MIT GenAI Divide Report explainer Why 95% of generative AI pilots left no profit footprint virtualizationreview.com

McKinsey State of AI 2025 key takeaways Why the 21% workflow redesign rate matters and how top performers pull 3x ahead gend.co

GitHub Copilot productivity RCT paper How 95 dev controlled trial proved 55.8% speed gain github.blog

BCG AI Impact Gap summary Only 25% of 1,800+ executives say they've realized measurable value tekstac.com

Klarna AI reversal timeline The full story from 700-job claim to rehiring digitalapplied.com

Three-step AI investment strategy for CFOs How to avoid FOMO and sunk-cost traps clobe.ai