The mainstream story about reinforcement learning in chip design is a story about scale. Bigger policy networks, more simulation, more compute thrown at the search space until something clicks. I think that story is mostly wrong, and the recent work on history-aware offline reinforcement learning for detailed routing is good evidence. A 92% reduction in routing violations alongside a 10% runtime cut did not come from a larger model or a longer training budget. It came from giving the agent a memory.
That distinction matters more than the headline number, and it generalizes well past electronic design automation.
Why Detailed Routing Breaks Naive Agents
Detailed routing is one of the least forgiving problems you can hand an agent. The task is to connect an enormous number of nets through a layout that is already congested, obeying design rules that interact in ways no single local decision can anticipate. Every wire you place consumes resources that later wires need. The agent’s early choices quietly determine whether the last 5% of nets are routable at all.
This is the structural reason standard reinforcement learning formulations struggle here. If you model routing as a Markov decision process where the state is the current grid occupancy, you are implicitly claiming the present configuration contains everything worth knowing. It does not. Two layouts can look identical in occupancy terms and yet sit in completely different positions relative to a violation, because how you arrived at that occupancy encodes information about congestion pressure, detour debt, and which regions are about to become unroutable.
An agent with no history is forced to re-derive that context from scratch at every step, and it cannot. So it makes locally reasonable decisions that accumulate into a violation-heavy result. Then the usual response is to scale the model, which mostly teaches the policy to memorize more local patterns rather than fixing the representational gap.
What the LSTM Is Actually Doing
The choice of an LSTM here reads as unfashionable at first glance. In 2026, reaching for a recurrent network instead of a transformer stack looks like a step backward. I would argue it is the correct inductive bias for this problem, and not a compromise.
Routing is a sequential resource-consumption process with a strong recency structure. What happened in the last few hundred steps in the local neighborhood matters enormously; what happened ten thousand steps ago in a distant region matters much less. A recurrent state that compresses trajectory history into a fixed vector is a natural fit. It gives the agent a running summary of pressure and commitment without paying quadratic attention costs over a step sequence that can grow very long.
Put differently, the LSTM is not there to be clever. It is there to convert a partially observable problem into something closer to fully observable from the policy’s point of view. That is the single highest-value intervention available in this class of task, and it is cheap compared to scaling.
The Offline Part Is the Quiet Win
The offline framing deserves as much attention as the history component. Online reinforcement learning against a real router is punishing. Each episode is expensive, the reward is sparse and delayed, and exploration in a space this constrained produces mostly garbage trajectories. Offline learning sidesteps that by training on trajectories that already exist, including the accumulated output of conventional routers.
This is where the 10% runtime improvement becomes interesting rather than incidental. A common pattern with learned components is that quality gains arrive with an inference tax, and the system ends up slower overall. Getting fewer violations and less runtime suggests the agent is not doing more work per decision. It is doing less wasted work: fewer detours, fewer rip-up-and-reroute cycles, fewer dead ends that have to be unwound. Better decisions early mean less cleanup later.
The Transferable Lesson for Agent Architecture
I read this result as a correction to a bias in how we build agents generally. When an agent underperforms, the reflex is to assume insufficient capacity or insufficient training. Often the real problem is that the state representation is hiding information the agent needs, and no amount of capacity recovers information that was never provided.
The diagnostic question is simple. Could an expert human make a good decision from exactly the observation your agent receives? For stateless routing agents the honest answer is no. A human engineer would immediately ask what has already been placed and where the pressure is building.
Adjacent work on photonic spiking reinforcement learning for routing points toward a different axis of progress, pushing the substrate rather than the representation. Both directions are worth pursuing. But the history-aware result is the one I would study first, because it is the cheaper lesson and the more portable one. Before scaling your agent, check whether it can remember what it just did.
🕒 Published: