Every commercial airliner carries more thrust than it needs at cruise. The engines are sized for the worst three minutes of the flight — takeoff on a hot day with a full load — and then spend the next six hours loafing. Electrical grids are built the same way: provisioned for the single worst hour of the worst day of the year, and underused the rest of the time. Data centers have historically behaved like an aircraft that insists on full takeoff power for the entire journey.
That assumption is what the Santa Clara pilot pokes at. In April 2026, Silicon Valley Power — the municipally owned utility serving the City of Santa Clara — and Emerald AI announced a program to demonstrate flexible data centers, with the goal of unlocking power capacity for AI without sacrificing performance. NVIDIA has backed Emerald AI, and the company’s work includes integration with NVIDIA DSX Flex. The mechanism is software that adjusts data center power consumption in response to grid conditions.
Stated that plainly, it sounds like a facilities story. I’d argue it’s an agent architecture story that happens to be wearing a hard hat.
Power becomes a scheduled resource
Any scheduler is a policy operating over a resource model. Kubernetes reasons about CPU and memory. GPU orchestrators add device memory, interconnect topology, and NUMA locality. What almost none of them model is the wall — the assumption is that electricity is infinite, instantaneous, and identical at every moment.
Once power draw becomes something a controller can modulate against external signals, the resource model gains a new axis, and it’s an unusual one:
- It is exogenous. Grid conditions are not set by your workload. The constraint arrives from outside the cluster, on someone else’s clock.
- It is temporal, not just quantitative. A megawatt at 3 a.m. and a megawatt at 6 p.m. are different goods. Schedulers are bad at pricing time.
- It is shared with parties who never signed your SLA. The neighborhood is a stakeholder in your training run.
Adding an axis like that to a scheduler is not a configuration change. It restructures what the objective function is even trying to optimize.
Two control loops that disagree
The interesting technical tension here is a clash of timescales. Grid operations run on cycles that stretch from seconds to hours. Inference serving runs on milliseconds, with tail latency as the thing users actually feel. Training runs on days, with checkpointing intervals that make interruption expensive but survivable.
A controller mediating between those loops has to be honest about which workloads are compressible and which are not. Batch training, fine-tuning jobs, embedding generation, offline evaluation, and speculative pre-computation are elastic. A synchronous agent trajectory with a human waiting at the end of it is not. The engineering work is classification and prediction: knowing what can yield, by how much, for how long, and what it costs to resume.
This is the same problem the agent community keeps rediscovering under different names. Deciding which sub-tasks to defer, which to run cheaply, and which demand full capability is the core of every planner-executor design. Emerald AI’s software solves that problem at the level of megawatts instead of tokens, but the shape of the decision is familiar — a policy with hard constraints on one side and a soft optimization target on the other.
Why the utility partner matters more than the vendor
Plenty of companies have built power-aware scheduling as an internal cost optimization. What distinguishes this pilot is the counterparty. A municipal utility participating directly means the flexibility has a recognized value on the grid side, not just a line item on an electricity bill. That turns a private efficiency trick into something closer to a two-way interface.
For those of us designing agent systems, that interface is the part worth watching. If power capacity becomes contingent on demonstrated flexibility, then flexibility becomes an architectural requirement that propagates upward — into how jobs are queued, how models are selected, how aggressively agents are allowed to fan out parallel tool calls, and how gracefully a system degrades when the answer is “not right now.”
What this asks of system designers
Most agent frameworks today assume capacity is a given and treat resource limits as an error condition. A rate limit is an exception to catch, not a signal to plan around. Systems built for a flexible power regime will need the opposite posture: degradation as a first-class path, with explicit policies for what gets sacrificed first.
That means quality-of-service tiers that mean something, checkpointing that is cheap enough to use often, and honest cost models for restart. It also means agent designs that can reason about their own budget — choosing a smaller model, a shallower search, or a deferred execution window when the environment says so.
The Santa Clara pilot is one program in one city, and it will be judged on whether it actually unlocks capacity while holding performance. But the premise underneath it — that compute can negotiate with the physical world instead of merely consuming it — is a genuinely different way to think about where a scheduler’s boundaries lie. Aircraft engines throttle back at cruise. It’s about time our clusters learned to.
🕒 Published: