From Token Maxing to Token Axing: Why Driving AI Adoption Instead of Outcomes Is Going to Get Expensive
· 5 min read · Nic Keating
Adoption dashboards measure activity, not results — and as AI shifts from chat to agents and the provider subsidies end, that blind spot turns into a cost curve you can't defend. Here's how to manage for outcomes per dollar before the next invoice lands.
Your AI adoption numbers are up. That might be the problem.
Most NZ executives I work with are running the same play. Buy the licences, drive adoption, report usage to the board. The logic is sound — nobody gets value from a tool they don't use. But the play has a trap in it now, and the trap is about to get expensive.
What Your Dashboard Actually Measures
Adoption metrics measure activity: licences activated, prompts per week, the percentage of staff using Copilot. None of them measures results. So the standard play builds a workforce that's used to AI without telling you which uses pay for themselves and which just burn money. That was fine while AI was cheap. The cheap part is ending.
Why the Bill Is About to Jump
The first shift is from chat to agents — software that works through a task on its own. The cost profile is nothing like chat. At every step, an agent re-reads everything it has already seen: its instructions, the files, the whole job so far. Across measured agentic workloads, that re-reading pushes input tokens to roughly 25 times output tokens.
GitHub put a number on it in April when it paused new Copilot Pro signups: a single agentic request can now consume over 500,000 tokens, costing more in compute than the entire monthly subscription.
The second shift is the end of the subsidy. For two years Microsoft, OpenAI, Anthropic and Google priced aggressively to win market share. Publicly tracked power users on Anthropic's US$200-a-month plan were burning through more than US$35,000 of compute a month, and the provider wore the difference. Nobody wears a 175-times subsidy forever.
GitHub moved every Copilot plan onto usage-based billing on 1 June. And on 9 June, Anthropic launched Claude Fable 5 — its most capable publicly available model yet — at US$10 per million input tokens and US$50 per million output tokens. Exactly double Opus 4.8, the model most enterprises are running today.
When the people selling the computing start rationing it, that tells you what it really costs them.
The Adoption Curse
Put those together. You've spent a year getting people used to AI, and that's worth having — the habits are real capability. But those habits are about to be billed by the token rather than by the seat, and your reporting can't tell you which of that use earns its keep. You end up funding a growing bill on a workforce that now expects the tool, with no way to defend the spend. The adoption curve you're celebrating is your cost curve, running about a year ahead.
Only One Response Ages Well
I'm watching two reactions across NZ this quarter. One group is building discipline: working out which use cases spend what, capping budgets per workflow, routing easy work to cheap small models and reserving the expensive ones for problems that earn them. Their usage isn't falling. Their cost per result is. The other group is panicking — switch off Copilot for half the workforce and wait. I've seen it in three organisations in the past month. A year spent running AI under a hard freeze teaches your people nothing except that the capability can be taken away.
What the Caps Taught Me
I learned this on my own bill before I saw it on anyone else's. I spend my working day inside these tools, and even after I'd optimised hard, the goalposts kept moving. A model I relied on would hit a weekly cap, or its pricing would shift, and work I'd tuned around it would stall. The fix wasn't a cheaper model. It was refusing to be locked to any one of them. I pulled my processes and context out of any single vendor's setup into a portable layer I own, so I can move the same work between models and vendors as the terms change.
A Word on Running AI on Your Own Hardware
One response I hear regularly from technology leaders: we'll push AI to the edge and run it on the laptops. It's a reasonable instinct. Local models have no metered cost and no vendor dependency. But most corporate estates in New Zealand aren't ready, and the refresh cycle problem is making that worse. I work across four corporate-issued laptops from different clients.
None can run a capable model locally; several are over eight years old. With PC prices rising and IT budgets flat, hardware refresh cycles are stretching out, not shortening. A strategy built on local inference your IT team can't execute for three to five years isn't a plan. It's a deferral. Portable architecture — keeping your data and processes in a layer you own — is the right goal. Local hardware is one way to get there eventually. Multi-vendor portability in the cloud is what's available now.
Four Moves Before the Next Invoice
Break the bill apart. AI spend usually lands as one line on a Microsoft, Google or AWS invoice. Split it by use case. Cost is almost always concentrated in a handful of workflows.
Right model, right job. Using a frontier model to summarise a meeting note is the most common waste I see. Easy work goes to cheap small models. Expensive models have to earn their place.
Keep yourself portable. Don't hard-wire your business to one provider. Hold your data and key processes in a layer you control, so you can shift to a different model — or vendor — when the terms change. Local hardware is the long-term play; multi-vendor portability is what's available now.
Measure outcomes per dollar. Decisions made, or revenue influenced, per dollar of AI spend. Prompts per week is the metric that got us here.
The Test
At your next AI review, ask one question: are we paying for outcomes, or paying the model to re-read its own notes? If the room goes quiet, you're not managing a capability. You're funding a habit, and the invoice is already in the post.