The most interesting thing about Jensen Huang’s claim that Nvidia can grow revenue 70% next year is that it is not a bullish statement. It is a supply statement dressed up as a growth statement, and that distinction matters more than the headline number.
The mainstream reading goes something like this: a company expected to close its current fiscal year near $400 billion in revenue projects roughly $680 billion, and that arithmetic either proves the AI trade is real or proves it has lost contact with reality. Speaking Thursday at Goldman Sachs’ Communacopia + Technology Conference, Huang repeated the figure — “I think we could grow 70% year over year. We’re confident about that” — and pointed to AI market dominance and demand that continues to exceed what the company can ship.
Note the shape of that reasoning. He is not forecasting that customers will want more. He is forecasting how much he can build. Those are entirely different kinds of prediction, and confusing them is how most commentary on this number goes wrong.
Demand forecasts are guesses, capacity forecasts are schedules
When a company says demand will grow 70%, it is making a claim about other people’s behavior. When a company says it is supply constrained and expects to grow 70%, it is making a claim about its own manufacturing pipeline, its packaging capacity, its memory allocations, and its order book. The first is speculation. The second is closer to logistics.
This is why the number feels absurd to people modeling AI as a consumer product cycle and feels almost mechanical to people who have watched inference bills at close range. If the constraint is on the supply side, the growth rate is a function of how fast you can add capacity, not how fast enthusiasm spreads.
Which raises the question I actually care about as someone who works on agent systems: what kind of workload produces demand that outruns the largest compute supplier on the planet?
Agents changed the shape of the compute curve
For most of the past few years, the mental model of AI compute was training-dominated. You spend an enormous amount of compute once, then serve the result cheaply. Under that model, a demand curve should eventually flatten. Training runs get bigger, but they are discrete events, and there are only so many organizations that can fund them.
Agentic systems break that model in a specific way. A single user request no longer maps to a single forward pass. It maps to a loop, and the loop has properties that are hostile to cheap serving:
- Multi-step reasoning means one request becomes many model invocations, and the count is data-dependent rather than fixed.
- Tool use inserts retrieved content back into context, so the sequence length grows as the task proceeds instead of staying constant.
- Verification and self-correction spend compute specifically to avoid wrong answers, which means quality improvements now cost tokens rather than saving them.
- Long-horizon tasks hold state across many turns, so memory pressure per active session climbs.
The economic consequence is that inference stops behaving like a marginal cost you optimize away and starts behaving like a capital input you buy more of. In a training-dominated world, better models reduce compute per unit of value. In an agent-dominated world, better models tend to raise it, because a more capable agent is trusted with longer, more expensive tasks.
What would actually falsify the forecast
I want to be careful here. Nothing in Huang’s public reasoning specifies workload mix, and I am not going to pretend otherwise. What he offered was market position and demand exceeding supply. The agent argument is my inference about why that demand persists, not something he said.
But it does suggest where to watch. If the 70% figure is really a capacity schedule, the thing that breaks it is not a shift in sentiment. It is a break in the physical chain — packaging, memory supply, power availability at the sites where these systems land. Sentiment can swing hard without changing a single delivery date.
The other thing that could break it is an architectural change on the software side. If someone finds a way to get agent-grade reliability with dramatically fewer model calls — better planning, better caching of intermediate reasoning, smaller specialized models handling most steps — then the demand curve bends for real reasons rather than emotional ones. That is a research question, not a market question, and it is the one I would fund if I had the choice.
Why the framing matters for builders
If you are designing agent architectures right now, the useful takeaway is not whether Nvidia hits $680 billion. It is that the entity with the best view of aggregate AI compute consumption is planning around demand it cannot fill. That is a signal about your own cost curve.
Systems that treat inference as nearly free will get expensive faster than their designers expect. Systems that treat each model call as a budgeted resource — with explicit stopping conditions, cheap paths for easy cases, and escalation only when needed — will look conservative for a while and then look correct. Compute scarcity tends to reward the engineers who assumed it.
🕒 Published: