\n\n\n\n Tokens Per Second Is the New Teraflop - AgntAI Tokens Per Second Is the New Teraflop - AgntAI \n

Tokens Per Second Is the New Teraflop

📖 5 min read•868 words•Updated Sep 25, 2026

It’s 11:40 on a Tuesday night and I’m watching a trace scroll by in a terminal. An agent is doing something unglamorous: reading a spec, deciding which of four tools to call, calling one, reading the result, deciding again. Twenty-two steps before it produces anything a human would call output. The model itself is fine. The reasoning is fine. What’s not fine is the clock in the corner of the screen, which has ticked past ninety seconds, and the fact that nineteen of those twenty-two steps were spent waiting on a single number I don’t control: how fast the inference layer can hand back the next token.

That number is why the AMD and Cerebras announcement matters more than a typical chip press release, and why Wall Street’s enthusiasm for the AI silicon names has a technical foundation underneath the momentum.

Why Latency Became the Architecture Constraint

For most of the past few years, the question we asked about inference hardware was throughput: how many requests per second, how many concurrent users, what does it cost per million tokens. That framing came from a world where the dominant workload was a chatbot. One human types, one model answers, the human reads for thirty seconds. Latency budgets were generous because the bottleneck was a person.

Agents broke that assumption. In an agentic system, the consumer of a model’s output is another invocation of the same model. There’s no human reading for thirty seconds. Every millisecond of time-to-first-token and every millisecond between tokens gets multiplied by the depth of the reasoning chain. A system that feels snappy in a single-turn benchmark can feel unusable once you wrap a planner, a critic, and a tool-execution loop around it.

This changes what “good hardware” means. Architectures optimized for batched throughput assume you can amortize cost across many simultaneous requests. Agent workloads are spikier and more sequential: a burst of dependent calls, then a pause while a tool runs, then another burst. The batching tricks that make throughput economics work fight against the latency profile agents need.

What AMD and Cerebras Are Actually Signaling

The joint announcement frames the offering as an ultra-low-latency, high-throughput inference solution. Notice that it’s both words, joined. That’s not marketing redundancy, it’s a statement about which constraint set the market now considers mandatory. Serving fast used to be a feature. Serving fast while serving a lot is now the table stakes for anyone who wants to host agent traffic.

Cerebras is interesting here precisely because its design philosophy diverges from the GPU lineage. A wafer-scale approach changes where memory sits relative to compute, and memory movement is where a lot of inference latency actually goes. Pairing that with AMD’s platform, CPU line, and ecosystem reach is a bet that the inference tier needs different silicon assumptions than the training tier. That’s a structural claim, not an incremental one.

AMD’s broader roadmap points the same direction. The MI400-based Helios server is slated for 2026 and positioned against Nvidia’s dominance, with Lisa Su emphasizing open collaboration. OpenAI has moved to adopt AMD’s newest chips. The 2nm EPYC work matters too, because agent systems are not pure GPU workloads. Orchestration, retrieval, vector search, tool sandboxes, and the glue code that makes a multi-agent system function all run on general-purpose cores. If your CPU tier is slow, your agent is slow no matter how fast the accelerator is.

The Facilitator Framing

There’s a reason both Nvidia and AMD want partners in this layer rather than trying to own every inch of it. The inference tier is turning into an integration problem. You need the accelerator, the host processor, the interconnect, the serving stack, the scheduler that understands dependent request chains, and the observability to figure out where the 400 milliseconds went. Nobody wins that alone, and the companies that make the pieces fit together are capturing value that used to sit entirely with whoever shipped the fastest FLOPs.

Wall Street’s optimism about AMD and Nvidia is often read as a bet on scale, more data centers, more chips, bigger clusters. I’d read the AMD-Cerebras move as a bet on something narrower and more interesting: that the shape of demand is shifting from training-heavy to inference-heavy, and within inference, from throughput-optimized to latency-optimized.

What I’d Watch as a Practitioner

  • Inter-token latency under realistic dependent-chain loads, not single-shot benchmarks
  • How the serving stack handles the bursty, sequential request pattern agents produce
  • Whether the CPU and accelerator tiers are designed together or bolted together
  • Portability, since being locked to one vendor’s inference characteristics constrains how you design agent loops

The uncomfortable truth for anyone building agent systems is that our architectural choices are downstream of hardware we don’t control. We design shallower reasoning chains because deep ones are too slow. We cache aggressively because calls are expensive. We avoid reflection loops that would improve output quality because users won’t wait. Those are not intellectual decisions, they’re latency decisions in disguise.

If the inference tier genuinely gets faster, some of those compromises stop being necessary, and the agent architectures we consider reasonable will look different a year from now. That’s the part of this story I’m watching, and it’s not the one showing up in the price targets.

đź•’ Published:

🧬
Written by Jake Chen

Deep tech researcher specializing in LLM architectures, agent reasoning, and autonomous systems. MS in Computer Science.

Learn more →
Browse Topics: AI/ML | Applications | Architecture | Machine Learning | Operations
Scroll to Top