Imagine building the only power plant in a boom town. Everyone comes to you for electricity, you set the rates, and your margins are the envy of the region. Then one of your largest customers, the one with the deepest pockets and the most land, quietly starts generating its own. Worse, it starts selling the surplus to your other customers.
That is roughly the shape of Nvidia’s current discomfort. The stock has stalled in its latest attempt at new highs, and Google’s accelerator business is a meaningful part of the reason. Nvidia’s fiscal 2026 numbers were not the problem: revenue of $215.94 billion, up 65.47% from $130.50 billion, with earnings of $120.07 billion. Those are the financials of a company executing extraordinarily well. The market is not pricing the past year. It is pricing the structure of the next several.
Why an Agent Researcher Cares About Silicon
I spend most of my time on agent architecture rather than semiconductors, so let me explain why this matters from where I sit. The economics of agentic systems are unlike the economics of a chatbot. A single agent run is not one forward pass. It is a loop: plan, call a tool, observe, revise, call again. Multiply by retries, by parallel branches, by verification passes that check the agent’s own work, and inference cost per completed task climbs by an order of magnitude or more compared to a one-shot response.
That changes what hardware buyers optimize for. Training runs are lumpy, enormous, and tolerant of exotic complexity, because you do them occasionally and you write custom kernels to squeeze every flop. Agent inference is the opposite: continuous, latency-sensitive, and enormous in aggregate. When your workload is a steady river rather than a seasonal flood, the calculus shifts from “what is the fastest chip” to “what is the cheapest sustained token, including power, interconnect, and the engineers I need to keep it running.”
Custom silicon is built for exactly that second question. A vertically integrated operator that runs its own models on its own accelerators in its own data centers captures margin at every layer. It does not pay a hardware vendor’s premium, and it can co-design the model architecture and the chip to suit one another. For workloads with a narrow, well-understood shape, which is what production agent serving eventually becomes, that integration is difficult to beat on cost.
The Shift From Captive Buyer to Competitor
Google building accelerators for internal use was always a cost story, not a competitive one. Nvidia lost some volume and kept the ecosystem. Selling those accelerators to outside customers is a different category of event. It converts an internal efficiency program into a product line, which means the comparison shoppers are no longer just Google’s own infrastructure teams. They are everyone else’s.
The reason this is a conundrum rather than a crisis is that Nvidia’s real product has never been only the die. It is CUDA, the kernel libraries, the years of accumulated tooling, and the fact that nearly every framework, profiler, and inference server in the field assumes Nvidia hardware underneath. Switching costs in that stack are not a line item. They are a staffing plan.
But software moats erode from the top down. The largest AI operators are precisely the ones who can afford to port their stack, because they employ people who write kernels for a living and because a few percentage points of inference cost across a fleet pays for that team many times over. Smaller shops stay on Nvidia because the alternative is not economical for them. So the competitive pressure arrives first at the highest-volume, highest-visibility accounts, which is where the growth narrative lives.
What I Would Watch
Analysts remain positive on Nvidia’s long-term growth, and the financials support that view. My read is that the argument is no longer about whether demand for compute keeps expanding. It clearly does, and agentic workloads are a structural reason why. The argument is about who captures the margin on that expansion.
- Whether serving frameworks and compilers become genuinely hardware-agnostic, which would turn accelerators into a commodity comparison rather than an ecosystem commitment.
- Whether agent workloads standardize enough in shape that fixed-function silicon can serve them without leaving performance on the table.
- Whether buyers split their fleets by task, training and experimentation on one vendor, high-volume inference on another, which is the most likely near-term outcome and the most quietly corrosive to pricing power.
Nvidia is not being displaced. It is being negotiated with, which is a new experience for a company that has spent this cycle as the only serious option. A stock that pauses on news about a competitor’s chip business is a stock whose valuation assumed no such business would materialize. The company is still growing quickly and still owns the default software stack. It just no longer owns the entire conversation, and for a position priced on structural inevitability, that distinction is the whole story.
🕒 Published: