Strategy Blog
The Inference Flip: Agents Are Now the Majority
Aaron Levie says agents already make up the majority of inference. SCX.ai sees the same shift in its own traffic. David Keane on what it means for companies running agents and the startups building them.
Aaron Levie, Box's CEO, posted something this week that deserves more attention than it got. Agents already make up the majority of inference, he wrote, and over the next year or two that will trend toward nearly all of it. The vast majority of tokens used in the world will be agents executing tasks for us in the background, 24/7.
I can confirm the direction of travel from our own traffic. At SCX.ai, agents are already driving a large and growing share of the inference capacity we serve. The pattern behind our busiest endpoints is not humans asking questions. It is programs calling programs, in loops, all night, every night.
Call it the inference flip. The customer for compute changed, and almost everything downstream of that changes with it.
What Agentic Traffic Looks Like
Human chat and agentic inference are different workloads wearing the same API.
A conversation is a short burst. A few hundred tokens in, a few hundred out, then silence while the human thinks. It follows the working day and the working week.
An agent is a loop. It plans, calls tools, reads results, corrects itself, and tries again. Every step re-feeds its own context, so the input grows while the task runs. It does not sleep, does not take holidays, and does not wait for anyone to type.
Microsoft Research published the first systematic study of this in April. Across eight frontier models on SWE-bench Verified, agentic coding tasks consumed on the order of 1000x more tokens than code reasoning or chat. Two findings mattered even more than the headline number. Input tokens, not output tokens, drive the overall cost. And the same task can vary by up to 30x in token usage between runs.
A conversation costs hundreds of tokens. A task costs millions. That is not a price difference. That is a different business.
What It Means if You Run Agents
If you are an enterprise deploying agents, inference has quietly changed shape on you. It is no longer a per-question cost that scales with employee curiosity. It is an operating line item that scales with how much work you hand to machines.
That makes cost per task a strategic metric, not a procurement detail. When an agent burns a million tokens to close a ticket, every fraction of a cent per token compounds across millions of tasks. The cheapest GPU is not the answer. The most efficient path from task to done is. Model choice, routing, and where your capacity actually sits decide that path before anyone writes a prompt.
There is a second dimension: what your agents touch. Levie's list is instructive. Agents deployed to read every code change to secure our software. Agents processing all of our data inside workflows. Agents doing the research behind recruiting and customer prospecting. Those are the crown jewels, running unattended. Where those agents run is now a governance question, not just a cost question. If you cannot answer for the infrastructure, you cannot really answer for the agent.
What It Means if You Build Agents
If you are a startup creating agents, the inference flip is your margin story.
Every agent company is an inference business whether it admits it or not. You sell outcomes priced in dollars, and you buy tokens priced in fractions of cents. The gap between those two numbers is your gross margin, and it is set at the inference layer: which models you route to, how well you batch and cache, and where the capacity sits.
That 30x run-to-run variance in token consumption is the hard part. You cannot price an agent on vibes. You need visibility into token spend per task, and you need infrastructure where the unit economics actually work at scale. The winners of the next wave of agent startups will treat inference as a first-class engineering and commercial decision from day one.
For Australian founders there is a bonus. You can build a world-class agent company here without shipping your margin, or your customer data, to infrastructure you do not control on another continent. That capability exists locally now. It did not two years ago.
Why We Built SCX.ai for This Wave
Levie closed his post with a line I agree with completely: incredible time to be doing anything in inference.
We built SCX.ai's sovereign inference capacity in Australia for exactly this demand curve. Not the bursty human chat of the last two years, but the sustained, input-heavy, 24/7 loops that agents run. Efficient serving, transparent token economics, and infrastructure our customers can point to when a regulator asks where their agents do their work.
Agents are the majority of inference now. The trend toward nearly all of it is already visible in the traffic we serve every night. The companies that run agents will win on cost per task and governance. The startups that build them will win on margin and speed. Both races are decided at the inference layer.
That is the layer we are building, here, for exactly this moment.
Sources
- Aaron Levie on X: "Agents already make up the majority of inference..."
- Microsoft Research: How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks
David Keane is the Founder and CEO of SCX.ai, Australia's sovereign AI infrastructure company.