The invoice you can't predict
Nobody can tell you what a token costs. That's not your failure, it's the business model. Token pricing scales with usage, not with the value you get back, and the venture subsidies masking the real cost of compute won't last. Here's what to shift before they run out.
If a great demo has ever made you reach for a budget you didn't have, you are not alone. An orchestrating agent managing a swarm of sub-agents. A workflow drafted in seconds. A CxO task done by a model before lunch. In the sandbox, it looks like magic. In production, the P&L tells a different story.
The problem is not the technology. It is how you are billed for it.
Token pricing has nothing to do with value
When you rent a closed-model API, you pay per token. Costs scale linearly with usage. Your business outcomes do not. A workflow that saves your team ten hours and a workflow that saves nothing cost the same per token. The only number you can reliably measure is the size of the invoice.
That mismatch is currently being masked. Several commercial labs price their models in ways that do not reflect the underlying cost of compute, financed instead by venture capital and venture debt. That subsidy will not last forever. When it ends, the invoice catches up to the real cost, and your margin absorbs the difference.
Ask your vendor what a token actually costs to produce. If the answer is unclear, budget accordingly.
The models are catching up. The pricing gap is not.
The raw models themselves are becoming commodities. The performance gap between closed frontier models and open-weight models, the kind you can download, run, and modify yourself, has narrowed to a matter of months rather than generations.
The cost gap has not narrowed at all. A complex multi-step workload can run into the dozens of euros per execution on a closed API. Run on open-weight models, the same workload costs a fraction of that in compute. For routine, high-volume enterprise tasks, defaulting to closed APIs is an old habit dressed up as a strategy.
Sovereignty is the forcing function, not the pitch
Cost is why this matters to your finance team. Jurisdiction is why it matters to your legal one. Data processed by a non-EU provider stays exposed to non-EU law, regardless of which region the server sits in. The CLOUD Act, DORA, NIS2, and the AI Act are not abstractions. They are the reason procurement is already asking questions your current vendor cannot answer.
Shifting high-volume workloads to open-weight models running on infrastructure you can actually inspect keeps your margins intact, keeps your data under EU jurisdiction, and removes the vendor lock-in that comes with a closed API you cannot audit, move, or price in advance.
What this looks like in practice
Pay-as-you-go, no upfront cost, no minimum commitment: you pay for the compute you use, not for a forecast someone else made about your usage. Look for infrastructure that publishes its prices instead of negotiating them behind a sales call. Look for a provider that can show you, not tell you, where the compute runs and on what.
That is also where the sustainability case belongs, third, not first. The energy behind your compute matters. But it is proof that a provider means what it says, not the reason to switch. Ask about the audit before you ask about the marketing.
The shift is already happening
Your competitors have started moving high-volume workloads off closed APIs and onto open-weight models running on sovereign infrastructure. The ones who wait are the ones who will explain the margin gap to their board when the subsidies run out.
We don't publish on a schedule. We publish when there's something worth your inbox.
Get the next one.
No weekly filler, no vendor hype. Just the numbers, the shifts, and the evidence, when we have them.