\n\n\n\n Everyone Names the GPU and Nobody Names the Substrate - AgntAI Everyone Names the GPU and Nobody Names the Substrate - AgntAI \n

Everyone Names the GPU and Nobody Names the Substrate

📖 4 min read•777 words•Updated Sep 17, 2026

A 2026 survey of ABF substrates in data center silicon makes a quiet observation that I keep coming back to: to deliver the compute and memory bandwidth frontier models demand, designers now place multiple large logic dies alongside stacks of memory on a single package. Not one chip. A neighborhood of them, sitting on a slab of resin-coated copper that almost nobody in the AI discourse can name.

That framing deserves more attention than it gets. When we say “NVIDIA chip” or “MI300X+,” we are naming one tenant and ignoring the building. And for agent workloads specifically, the tenants we ignore are the ones setting the rent.

What the accelerator actually is

Strip the marketing off a modern data center accelerator and you find three things that matter roughly equally: logic dies that do arithmetic, high-bandwidth memory stacks that feed them, and the package substrate that wires the whole assembly together at densities that were exotic a few years ago. The transistor count gets the headline. Meta’s third-generation MTIA, built on TSMC’s N3P with very likely more than 100 billion transistors and HBM memory, will be discussed almost entirely in terms of that first number.

But the number that determines whether your agent finishes a 40-step tool-use loop in two seconds or twelve is bandwidth, and bandwidth lives in the memory stack and the interconnect. This is not a subtle secondary effect. It is the primary constraint for the class of work that agent systems generate.

Agents are a memory problem wearing a compute costume

Training is compute-bound in the way the industry’s mental model assumes: dense matrix math, big batches, high arithmetic intensity. Agent inference is not that. An agent doing planning, retrieval, and tool calls produces:

  • Long, growing contexts that inflate KV cache far beyond model weights
  • Serial decode steps where each token depends on the last, so you cannot batch your way out of latency
  • Frequent context switches across many concurrent sessions, each with its own cache state
  • Short bursts of generation between long waits on external tools

Every item on that list stresses memory capacity, memory bandwidth, and the cost of moving state around. None of them are solved by adding FLOPS. An accelerator with enormous arithmetic throughput and insufficient bandwidth will sit mostly idle during agent decode, burning power to wait.

The money is already voting

If you want evidence that the field understands this, look at where capital went in early 2026 rather than at conference keynotes. Positron AI raised $230M in a Series B on February 4, 2026, for a memory-centric inference accelerator. The description is the thesis. On February 3, one day earlier, Cerebras closed $1B in a Series H for its wafer-scale training and inference processor.

Cerebras deserves note here because the Wafer Scale Engine 3 is, in effect, an argument about packaging. If the substrate and the off-die interconnect are where your latency and energy go, one response is to stop cutting the wafer into separate dies and stop paying the crossing cost at all. Whether that is the right answer commercially is a separate question. As an engineering position, it is a direct response to the same constraint Positron is attacking from the memory side.

Why the market share numbers are more interesting than they look

NVIDIA is estimated at roughly 80 to 85 percent of data center AI accelerator revenue in 2026, down from around 92 percent in 2024. AMD is up from roughly 2 percent to something in the 5 to 7 percent range. Those movements are modest in absolute terms and enormous in signal terms, because the erosion is not coming from a single frontal assault.

It is coming from several directions with different theories of the workload. AMD’s MI300X+ competes on memory configuration. Broadcom’s custom XPUs are gaining traction precisely because a hyperscaler that knows its own serving profile can specify memory and interconnect for that profile instead of buying a general-purpose part. Meta’s MTIA line is the same logic taken in-house. Meanwhile Intel’s Gaudi 3 underperforms, which is a useful negative result: a credible logic design is not sufficient if the surrounding system does not match how the work actually behaves.

What this means if you build agent systems

Practically, it means you should evaluate hardware the way you evaluate a database, not the way you evaluate a benchmark score. Ask about memory capacity per accelerator, bandwidth per unit of capacity, and what happens to your KV cache when concurrency triples. Ask what the cost is when state has to cross a package boundary, because for agent traffic it crosses constantly.

The specification sheet will keep leading with the logic die. The performance you experience will keep being decided by everything soldered around it.

🕒 Published:

🧬
Written by Jake Chen

Deep tech researcher specializing in LLM architectures, agent reasoning, and autonomous systems. MS in Computer Science.

Learn more →
Browse Topics: AI/ML | Applications | Architecture | Machine Learning | Operations
Scroll to Top