Klarna announced in early 2024 that its AI chatbot was doing the work of 700 people. They'd cut support staff and save $40 million a year. Then, somewhere between late 2025 and early 2026, they quietly started rehiring.
It wasn't the number itself that was wrong. The problem was that it answered a question the number couldn't actually answer. When you try to respond to "what's our AI ROI?" with a single multiplier across the whole company, this pattern repeats itself.
Why 700 jobs flipped in six months
Look back at Klarna's actual numbers and the answer's right there. The AI handled straightforward inquiries fine, but it kept hitting walls with complex payment disputes, fraud reports, and policy exceptions. Customers called back for the same issue multiple times. The chatbot struggled to calm upset customers. The company ended up going hybrid: AI handled the first response, humans dealt with exceptions and customer retention.
The real issue wasn't the "700 jobs" figure itself. It was that the figure never said which tasks produced it. On simple queries, maybe it really did equal 700 people's worth of work. The moment you scale that number across all customer service, though—where 20-30% of tasks need exceptions—the costs start leaking right back out.
Why '2.5x ROI' figures are dangerous
This isn't a Klarna-specific story. That's what makes it scary. Harvard Business School and BCG ran a field experiment with 758 consultants. On tasks within AI's strength zone (inside the boundary), AI delivered 25% faster results and 40% higher quality. Step slightly outside that boundary, and consultants using AI were 19 percentage points more likely to produce wrong answers. The team called this the "jagged frontier"—same workflow, some tasks where AI crushes it, the next task over where it's actually worse than humans.
The instant you flatten "our company's AI ROI is 2.5x" into one number, you're averaging together what the tool does well and what it doesn't. McKinsey found that 88% of organizations using generative AI say they're using it regularly. Only 21% actually redesigned workflows to make it work. That redesign gap? It's the biggest predictor of EBIT impact. MIT's research went harder: 95% of the generative AI pilots they studied didn't produce measurable profit improvement at all. BCG surveyed over 1,800 executives. 75% called AI a priority. 25% said they'd actually realized meaningful value.
| Company-wide flattened number | Task-level validation benchmark | |
|---|---|---|
| Typical example | "AI cuts 700 jobs" | Customer support, standardized inquiries -15% handling time |
| Source | Company press release | Peer-reviewed field experiment / controlled RCT |
| Repeatability | Breaks when exceptions mix in | Clear task boundaries make reproduction possible |
| Klarna's outcome | Reversed with rehiring six months later | Developer productivity +55.8% holds up |
So what can you actually tell your CFO?
Credible numbers exist. They just come from places with clear task boundaries. GitHub ran a controlled experiment with 95 pro developers split randomly. The Copilot group finished the same HTTP server build task 55.8% faster (1h 11m vs 2h 41m), with higher quality too (78% vs 70%). Statistically significant (p=0.0017).
These numbers share one trait: source credibility comes in tiers. A vendor's customer story (Klarna's 700 jobs) shouldn't sit in the same category as a randomized controlled trial (Copilot's 55.8%). Stack them by credibility and you get this.
| Credibility | Source type | Example |
|---|---|---|
| ① Highest | Peer-reviewed field experiment | HBS/BCG consultant experiment, GitHub Copilot RCT |
| ② High | Investor disclosure | Financial figures in earnings reports |
| ③ Medium | Internal operational case | Savings verified on internal dashboards |
| ④ Medium-low | Vendor customer story | "Our customer saved 700 jobs" |
| ⑤ Low | Single-executive survey | Executive gut-feel response (use only cross-checked with other sources) |
Klarna's number lands in tier ④—vendor customer story. That tier's for reference, not for board reports as-is. Copilot's 55.8% or the consultant experiment's 25%? That's tier ①, safe to cite much more directly.
Five steps to a defensible ROI matrix
- Break work into tasks
Don't lump "customer service AI rollout" into one bucket. Separate "standardized inquiries" from "payment disputes" and measure each. - Lock in baseline per task
Record volume, time, cost, quality, and exception rates before deploying AI. Without this, you have no way to prove later impact. - Tag source credibility
Use the five-tier table above. Mark tier ④ and ⑤ numbers as "reference only" and keep them separate in reports. - Don't mix forecast, annualized, and realized savings
"$40M projected" and "cost actually cut this quarter" go on different lines. Mix them and Klarna's story repeats. It's also the CFO's job to stop pilots from dragging on due to sunk-cost bias. - Note workflow redesign separately
Same tool, different outcomes between teams that redesigned process vs. teams that just plugged it into old workflows. Track redesign as its own column.
Watch out
When you expand AI beyond where it works well, be especially careful. Using success in one task to justify pushing into exception-handling and emotional work can drop accuracy below human performance.
Common questions
We've got a pilot already running. How do we validate its ROI after the fact?
Without a pre-deployment baseline, perfect validation's impossible. Instead, find a similar task in a team or location that hasn't adopted AI yet and run a quasi-experimental comparison. If that's not possible, at least set a fresh task-level baseline starting now and compare quarter-over-quarter going forward.
Every department measures differently. How do we get them comparable?
Don't force one global metric. Lock in just five items as department-wide standards: volume, time, cost, quality, exception rate. Keep everything else task-specific. When those five match, departments' multipliers may differ, but credibility tiers stay comparable.
Can we just report vendor-supplied numbers straight to the board?
If you tag them properly, yes. Call it "vendor-cited tier ④" and note it's not reproduced internally. That way, if it flips like Klarna did, your credibility doesn't crater.
Is there a minimum ROI unit we can measure without redesigning workflows?
Yes. Tasks with naturally clear boundaries—code review, standardized support responses—show repeatable results even plugged into old workflows. Tasks with mixed exception-handling don't hold up without redesign.
How do we spot failure signals early in a pilot?
Watch repeat-inquiry rates and handoff-to-human rates. Same question coming back multiple times, or AI not reducing the rate it hands off to staff? That task isn't inside AI's boundary yet.
Go deeper if you want
Navigating the Jagged Technological Frontier Original research with 758 consultants showing which tasks AI gets right and which it doesn't aiinstitute.hbs.edu
MIT GenAI Divide Report explainer Why 95% of generative AI pilots left no profit footprint virtualizationreview.com
McKinsey State of AI 2025 key takeaways Why the 21% workflow redesign rate matters and how top performers pull 3x ahead gend.co
GitHub Copilot productivity RCT paper How 95 dev controlled trial proved 55.8% speed gain github.blog
BCG AI Impact Gap summary Only 25% of 1,800+ executives say they've realized measurable value tekstac.com
Klarna AI reversal timeline The full story from 700-job claim to rehiring digitalapplied.com
Three-step AI investment strategy for CFOs How to avoid FOMO and sunk-cost traps clobe.ai




