When you read that an accelerator delivers 30x faster inference, what exactly do you picture doing the work? Most people picture the die. The rectangle of silicon with the logo on it, the thing that gets a codename and a keynote slide. That picture is incomplete in a way that matters more every year.
Cerebras claims its CS-4 hits 30x faster inference than GPUs. Google’s Ironwood has reportedly passed Nvidia’s Blackwell on performance. Both of those numbers are products of an entire assembly, not a single chip. Underneath the compute die sits a layer almost nobody outside the supply chain talks about: the package substrate, and specifically the ABF substrate that has quietly become one of the hardest parts of building data center silicon in 2026.
What a benchmark number actually contains
A performance figure on an AI accelerator is a composite. It reflects the logic die, yes, but also how many high bandwidth memory stacks sit next to that die, how short and dense the wiring between them is, how much heat the package can move before the clocks throttle, and how many of those assemblies survive manufacturing intact. The substrate is the floor all of that stands on. It routes thousands of connections between silicon and board, holds everything in mechanical alignment as the package heats and cools, and does it across a footprint that keeps growing because modern accelerators are less “a chip” and more “a neighborhood of chips.”
That is why ABF substrate supply keeps showing up in trade coverage. It is not a glamorous component. It is a build-up resin layer with extremely fine features, produced by a small set of suppliers, on equipment with long lead times. When it constrains, it constrains everyone at once, regardless of whose architecture is winning the benchmark that quarter.
Why this is an agent architecture problem
I care about this for reasons that have little to do with hardware fandom. The workloads I study have changed shape. A single chat completion is a fairly polite guest. An agent is not. An agent loops. It plans, calls a tool, reads a result, revises, calls again, and carries a growing context through every step. The compute per token may be modest, but the memory traffic and the round trips are relentless.
That shape stresses exactly the parts of the system the substrate governs. Keeping a long context resident and reachable is a memory capacity and bandwidth question, which means it is a question about how many memory stacks you can place next to the compute die and how well you can wire them. Coordinating many concurrent agent sessions is an interconnect question, which means it is a question about how many signals leave the package and at what quality. Sustaining that over hours of continuous agent activity, rather than bursts of chat traffic, is a thermal and mechanical question, which is once again the package.
So when Cerebras builds around a wafer-scale part and Google builds Ironwood for its own fleet, they are making different bets about where the hard limit sits. Neither bet is really about arithmetic throughput. Arithmetic has been cheap for a while. What is expensive is feeding it and holding it together.
Capital is arriving, on a slower clock
Money is moving into chip ventures at a healthy pace. OLIX raised $312 million in a Series B in 2026, one of several rounds that show investors are no longer treating silicon as too capital-heavy to touch. That is a real shift in appetite.
It also runs into a timing mismatch. Software iterates in weeks. Model architectures turn over in months. Substrate and advanced packaging capacity turns over in years, because it involves buildings, tooling, and process qualification. Funding a design team is fast. Adding a production line for fine-feature substrates is not. A round announced today does not change what can be shipped next quarter.
What I would rather read
Comparisons between accelerators would be far more useful with a few additional facts attached. How much memory sits in the package, and at what bandwidth. What the sustained figure looks like versus the peak, measured over a long-running workload rather than a short one. How the part performs on multi-turn agent traffic, not just single-pass generation. And how many units actually exist, because a part that cannot be packaged in volume is a research result, not a product.
None of that is as quotable as 30x. But the constraint on running agents at scale in 2026 is turning out to be less about who has the fastest die and more about who can reliably build, feed, and cool the assembly around it. The unnamed layer at the bottom of the stack has become one of the more interesting variables in the entire business.
đź•’ Published: