\n\n\n\n Analysts Are Bullish on Chips for the Wrong Reasons - AgntAI Analysts Are Bullish on Chips for the Wrong Reasons - AgntAI \n

Analysts Are Bullish on Chips for the Wrong Reasons

📖 4 min read•787 words•Updated Sep 23, 2026

Wall Street’s enthusiasm for Nvidia and Broadcom heading into 2026 is probably correct, and almost entirely mispriced in its reasoning. Bernstein’s analysts remain bullish on both names. Stacy Rasgon’s colleague Arya points to companies with “moats that are quantified by their margin structure.” Morningstar lists Broadcom among undervalued AI plays as of July 2026. The consensus argument reduces to a story about compute scarcity: models keep getting bigger, training runs keep getting more expensive, and whoever sells the shovels wins.

That story is stale. It describes 2023. The thing actually driving silicon demand now is not model size. It is agent architecture, and the difference matters enormously for which margins hold and which evaporate.

Training scaled compute. Agents scale coordination.

A monolithic model call is a clean workload. One forward pass, predictable memory access, batch it and move on. That workload rewards raw FLOPs, and it made the GPU the default unit of AI infrastructure.

An agent does something structurally different. It plans, calls a tool, waits, reads the result, revises, calls another tool, maintains state across dozens of turns, and often spawns sub-agents that do the same thing in parallel. The compute pattern stops looking like a matrix multiplication and starts looking like a distributed system with a language model sitting at each node.

The bottleneck shifts accordingly. In a single large training run, interconnect matters but compute dominates. In an agent swarm serving thousands of concurrent sessions, the expensive parts are the ones nobody puts on a keynote slide:

  • Key-value cache management across long, branching context histories
  • Memory bandwidth, not arithmetic throughput, as sessions hold state
  • Network latency between the model, the tool layer, and the orchestration logic
  • Tail latency, because an agent is only as fast as its slowest sequential step

Those are networking and memory problems dressed up as AI problems.

Why Broadcom’s position is stronger than its story

This is where I think the analyst consensus accidentally arrives at the right stocks. Broadcom’s strength is described as leadership in networking and custom silicon, which sounds like a sensible second-tier bet next to the GPU monopoly. Under an agent-heavy workload mix, it reads differently. If the dominant cost of running agents is moving state around rather than multiplying matrices, then switching fabric and custom accelerators built for a specific inference pattern stop being commodity components. They become the part of the stack that determines whether your agent responds in 400 milliseconds or four seconds.

Custom silicon also fits the economics of agent deployment better than general-purpose hardware does. A hyperscaler running one enormous frontier model wants maximum flexibility. A company running ten thousand narrow agents that each do roughly the same shape of work wants a chip tuned to that shape. That is an ASIC argument, and it favors whoever designs ASICs for people who know exactly what they need.

The moat is software, and software moats are workload-specific

The bull case for Nvidia leans hard on ecosystem lock-in. One commentator put it as an “unassailable monopoly” built on owning the software layer on top of the hardware advantage. I agree that CUDA is the real asset. I disagree that it is unassailable in the way people mean.

CUDA’s grip is strongest where the abstraction it provides is the one developers need. For kernel-level work on dense tensor operations, nothing comes close. Agent frameworks, though, live several layers up. The developer building a multi-step agent is reasoning about tool schemas, retry policies, context windows, and cost per task. They are not writing kernels. They are writing orchestration code against an inference endpoint, and that endpoint’s underlying hardware is increasingly an implementation detail they never see.

Lock-in that operates two abstraction layers below the developer’s actual concern is weaker lock-in. Not absent, weaker. That is the asymmetry I would be watching if I held these positions.

Where I land

The conclusion overlaps with Bernstein’s, reached by a different route. Nvidia and Broadcom both benefit, but from separate mechanisms. Nvidia captures value as long as frontier training and high-end inference remain concentrated and margin-rich. Broadcom captures value as agent deployment pushes the cost center toward interconnect and purpose-built silicon, which is the direction the technical evidence points.

The risk in the consensus view is that it treats these as the same bet on the same trend. They are not. If agent architectures come to dominate production AI spending, the networking and custom silicon side of the thesis gets stronger while the general-purpose compute side faces pressure from specialized alternatives it currently dismisses.

Margin structure as a moat, as Arya framed it, is a good test. It just needs to be applied to the workload that is actually growing, rather than the one that made everyone rich two years ago.

🕒 Published:

🧬
Written by Jake Chen

Deep tech researcher specializing in LLM architectures, agent reasoning, and autonomous systems. MS in Computer Science.

Learn more →
Browse Topics: AI/ML | Applications | Architecture | Machine Learning | Operations
Scroll to Top