Back to Blog

ai

Faster, Cheaper Tokens — Why Your AI Bill Won't Fall Like You Think

6 min read · · By Kiran Gawde

Diagram titled 'Unit price collapses. Total spend still climbs.' A line chart indexed to 2023 = 100 plots two series on one axis from 2023 to 2027: cost per token falling from 100 to 4 (a 96% decline per token) and total AI spend rising from 100 to 395 (a 295% increase). A side panel defines Jevons' paradox — make a unit cheaper and total consumption rises faster than price falls — then lists costs that ignore the curve (longer context, multimodal and agentic loops consuming more tokens per finished job; frontier reasoning staying premium as tiers bifurcate; data infrastructure, evaluation, orchestration, security and human review; agent sprawl turning a $0.01 task into a $1.00 task) alongside what holds the line (route by complexity, measure cost per successful outcome, bound the agent loops). A note marks the index as illustrative rather than vendor price data

The internet got dramatically faster. The price per megabit collapsed. And yet, no enterprise CFO looked at their connectivity budget in 2020 and said, “We’re spending less on networking than we did in 2005.” Total spend went up — because usage exploded to fill every efficiency gain and then some.

LLM inference is on the same curve. Unit costs are falling fast. Total AI bills are not going to follow.

If you’re a VP of Digital, Product, or Engineering setting AI platform budgets for the next 12–24 months, this distinction is the most important thing you can internalize right now.


The Bandwidth Analogy Is More Precise Than It Looks

Nielsen’s Law — articulated by Jakob Nielsen of NN/g — holds that user bandwidth grows at roughly 50% per year. The FCC’s own broadband performance research documents that advertised fixed broadband speeds grew at approximately 20% per year from the late 1990s onward, roughly doubling every four to five years. Prices did not rise in lockstep. The result: the cost per usable megabit fell sharply, decade over decade.

The directional trajectory of U.S. consumer download speeds tells the story clearly. Compilations from sources like Ooma and Allconnect point to speeds moving from sub-1 Mbps in the early 2000s, to roughly 10 Mbps around 2010, to tens of Mbps by the mid-2010s, to 100+ Mbps in the early 2020s, and toward 200 Mbps by the mid-2020s. Monthly household bills, by contrast, often stayed within a broadly similar nominal range. Orders-of-magnitude improvement in capability; no corresponding collapse in total spend.

Why? Because Cisco’s historical traffic data shows that global internet traffic has grown by a multi-billion-fold factor over the past three decades. More capacity enabled more video, more cloud, more always-on applications — none of which existed before the capacity was cheap enough to make them viable. The bandwidth market demonstrated what economists call Jevons’ paradox in near-textbook form: make a resource cheaper per unit, and total consumption of that resource rises, not falls.


LLM Inference Is Running the Same Playbook

The a16z “LLMflation” thesis makes the parallel explicit: inference cost for equivalent capability has been falling rapidly, tracking a curve similar to other technology cost declines. Researchers at Epoch AI and independent price trackers have documented steep multi-year declines in cost per million tokens across the industry from 2022 through 2026.

The illustrative numbers here require care. OpenAI’s GPT-4 API at launch in March 2023 was widely cited at approximately $30 per million input tokens and $60 per million output tokens. By 2025–2026, mid-tier and efficiency-class models — including later OpenAI pricing tiers, open-weight competition, and purpose-built inference providers — pushed equivalent-capability pricing down by significant multiples. Some secondary analyses have described declines on the order of 10× per year for a given quality tier, or 95%+ reductions over roughly two years. Treat those specific figures as directional and sourced from secondary trackers rather than primary provider documentation; verify against current provider price sheets before building financial models on them. The direction, however, is not in dispute.

The engineering drivers are also well understood: better accelerators and higher GPU utilization; quantization and model distillation reducing compute requirements per inference; speculative decoding and improved serving stacks; and intensifying competition from open-weight models that no single vendor controls. These aren’t one-time gains. They are compounding forces, and they will keep pushing the unit price down.


Where the Analogy Holds — and Where It Breaks

The bandwidth parallel is useful, but not perfect. Knowing where it diverges is where the real budget planning happens.

Where it holds:

  • The unit price of “capability bits” — what you get per dollar of inference — will keep falling.
  • As that happens, new applications that weren’t viable at 2023 prices become viable. Volume grows.
  • Volume grows faster than price falls. Total category spend rises. This is the Jevons dynamic again, and there is no reason to believe AI workloads are exempt from it.

Where it breaks:

The capability treadmill moves fast. Internet speeds improved, but a web page in 2005 was still a web page. LLM workloads don’t stay fixed. Longer context windows, multimodal inputs, and multi-step agentic workflows all consume dramatically more tokens and FLOPs per completed “job” than a simple chat completion. The task you’re automating in 2027 will be more complex than the task you’re automating today, and that complexity has a token cost.

Quality tiers are bifurcating, not converging. Commodity chat, autocomplete, and classification tasks are approaching very low cost per token. Frontier reasoning — complex multi-step analysis, code generation at the architecture level, judgment-intensive workflows — is staying premium. Buyers who model a single blended rate are going to be wrong in both directions.

TCO extends well beyond the API call. Once token prices get cheap enough to stop dominating the conversation, the costs that remain are data infrastructure, evaluation pipelines, integration and orchestration engineering, security and compliance overhead, and human review for high-stakes outputs. These costs don’t follow the inference price curve. In many mature AI deployments, they already dominate it.

Waiting for “free AI” is a compounding mistake. Every quarter an organization defers building AI-native workflows, its competitors are accumulating learning, data, and operational muscle that doesn’t transfer. The cost curve is not a reason to wait. It’s a reason to design carefully now so you’re positioned to capture the gains as they arrive.


What This Means for Your Budget and Architecture

Four principles that should shape how you plan AI platform spend right now:

1. Plan for declining unit cost AND rising total workload simultaneously. These are not contradictory. They are the expected outcome. Budget accordingly rather than treating one as canceling out the other.

2. Architect model routing from day one. High-volume, lower-complexity paths — content classification, retrieval augmentation, structured extraction — should route to small, efficient models. Frontier capability should be reserved for workflows where the incremental quality demonstrably justifies the cost. The organizations getting this right are measuring cost per workflow, not staring at a single API bill.

3. Measure cost per successful business outcome. Cost per million tokens is an infrastructure metric. It tells you almost nothing about whether your AI investment is working. Define the outcome — a resolved support ticket, a completed product configuration, a qualified lead scored correctly — and measure cost against that unit. This is what makes the model routing decision legible to finance and product leadership, not just engineering.

4. Don’t let agent sprawl swallow the efficiency gains. Agentic architectures can turn a $0.01 task into a $1.00 task through unnecessary tool calls, redundant context loading, and poor loop termination. The inference price curve gives you margin to work with. Undisciplined orchestration will consume it faster than the curve delivers it.


The Bottom Line

Cheaper tokens are coming. They are already here at scale compared to where the market was 30 months ago. But the lesson from three decades of internet infrastructure isn’t that cheap bandwidth lowered network budgets — it’s that cheap bandwidth enabled entirely new categories of consumption that expanded the budget envelope while transforming what organizations could build.

The same dynamic is playing out in AI. The opportunity is real. The cost discipline required to capture it without runaway spend is also real.

Object Edge works with digital and commerce organizations to design AI-native operating models and knowledge orchestration architectures that exploit the inference cost curve without unconstrained agent sprawl. If you’re building out your AI platform budget or rearchitecting your inference stack, let’s talk.

Related articles

Let's build something extraordinary

Ready to accelerate your digital transformation? Talk to our team.

Add Object Edge as a preferred source on Google ↗ to see our articles highlighted in AI Overviews and Top Stories.