A token is a physical product

A token isn't just text, it's silicon, real estate, and power converging into a physical product with a physical cost. Closed models can run 100x more expensive than open-weight alternatives for the same task, and carbon rarely makes the invoice at all. Here's the equation behind the price.

Share
A token is a physical product
Photo by Shubham Dhage / Unsplash
💡
Part 2 of 6 in a series on AI economics

What it physically takes to produce one token, and what you're actually paying for.

You've read the word a hundred times. A model generates words, code, or images, and it gets metered in tokens. In daily use, tokens feel weightless, invisible units of text passing through an API.

But on your P&L, and on the planet, a token is a physical product. It has a physical cost, a physical jurisdiction, and a physical carbon footprint. Understanding AI unit economics means looking behind the invoice and asking what it actually takes to produce one.

The physical bill of materials

Three things have to converge to produce a token.

  1. Silicon. High-performance GPU clusters that demand heavy capital, complex supply chains, and increasingly, geopolitical approval just to buy.
  2. Real estate. High-density server rooms built to house those chips and shed the heat they produce.
  3. Power. Energy to run the processors, then energy again to cool them.

None of this is free, and most data centres still draw from mixed-energy grids, at an average of over 400g CO₂ per kWh. Rent a closed API and you are quietly paying for wherever that grid happens to be.

The scaling trap, and the agentic multiplier

The old playbook, brute-force scaling, is hitting a wall. imec's research shows power consumption rising close to linearly with hardware: process ten times the workload, and you need roughly ten times the hardware and ten times the electricity.

Autonomous agents make that trap tighter. Because they call a model repeatedly to plan, execute, and self-correct, KAIST researchers measured agents consuming up to 136.5 times more energy per query than a single chatbot answer. Getting ahead of this gap will take real efficiency gains across algorithms, hardware, and energy systems together, not scaling one layer in isolation.

Closed versus open-weight: the markup is real

Then look at the price tag. A closed frontier model like Claude Fable 5 lists at roughly €50 per million output tokens.

Independent evaluations by the AI Safety Institute put the gap in concrete terms: running a 100-million-token workload on a top-tier closed model costs roughly €85. A comparable open-weight model, such as DeepSeek V4-Pro, run on first-party pricing, costs roughly €1.19. Close to a 70x difference for comparable work.

Is the premium ever worth it? On the hardest tasks, sometimes. But for routine, high-volume workloads, that markup is difficult to defend, which is why more teams are shifting those workloads to open-weight models on private compute. The question has moved from "which model is best" to "what do I keep control of."

The cost that rarely makes the invoice: carbon

One line almost never appears on the bill: carbon. If compute is a physical commodity, its waste has to be accounted for too. Most data centres draw energy from the grid, generate heat, and vent it straight into the atmosphere.

In Europe, CSRD makes scope 3 carbon tracking a legal requirement, not an optional one, and the first conflicts between citizens and data centre operators over energy priority have already started. Sustainability does not have to be an added cost here. Carbon-aware scheduling, running heavy workloads when clean energy is abundant, and circular heating, capturing server heat to warm nearby buildings, reduce the footprint and the cost per token at the same time.

Build your unit economics on the equation, not the invoice

A price is a decision someone else made. A cost is an equation you can check yourself. Three levers to check against your current vendor:

  1. Zero venture markup. Providers running shared GPU pools with usage-aware scheduling can offer open-weight models close to raw compute cost, without the closed-lab premium built in.
  2. Low carbon footprint. Compute run on local renewable energy, scheduled around when that energy peaks, cuts both emissions and price per token.
  3. Circular heating. Server rooms integrated into mixed-use buildings can turn waste heat into heating for the building next door, closing the carbon loop and the cost loop at once.

Do you actually know what your organisation pays per million tokens, and what that price is built on? If not, that number is worth getting before the next renewal. You can stop renting foreign, carbon-heavy pipelines at a markup you can't audit, and run your workload on infrastructure where the compute, the jurisdiction, and the energy source are all visible to you.

We don't publish on a schedule. We publish when there's something worth your inbox.

Get the next one.
No weekly filler, no vendor hype. Just the numbers, the shifts, and the evidence, when we have them.

Subscribe to the newsletter