From Megawatts to Tokens: How NVIDIA Maximizes AI Factory Production

On a sweltering August evening in Silicon Valley, as the sun dropped and air conditioning loads spiked , Silicon Valley Power sent a signal to an AI factory to adjust its power consumption.
Varun Sivaram was watching on Zoom with about forty others — his team at Emerald AI in their San Francisco conference room, engineers at the data center and people from the utility itself. Nobody touched anything.
Emerald AI’ s Conductor platform — a grid-orchestration platform from NVIDIA partner Emerald AI, and an early example of the kind of flexibility NVIDIA DSX Flex is built to deliver — receives signals about grid conditions and adjusts the data center’s flexible computing workloads. Work that can wait is slowed or rescheduled, while higher-priority services continue operating.
The goal is to reduce electricity demand when the grid is constrained without interrupting critical AI workloads— exactly what Silicon Valley Powe r needed,
“We were watching with bated breath,” Sivaram said. “It was our first time deploying across thousands of NVIDIA GPUs.” His head of product, Mansi Shah, was emotional. “This feels kind of like a SpaceX rocket launch,” she said.
Silicon Valley Power h as since sent more than 200 demand signals to that AI factory. It worked every single time.
This is grid flexibility in production.
And it points at something much bigger than one facility in Santa Clara: a path to unlocking the power America’s AI factories need, without waiting a decade to build new transmission lines.
At the AI Infra Summit on Tuesday, Ian Buck, NVIDIA’s vice president of hyperscale and high-performance computing, made AI factory efficiency the centerpiece of his infrastructure keynote.
Results from cloud provider Lambda’s first validation in a deployment environment, released the same day, put numbers to it: a fixed power budget can support 24% more token throughput when managed intelligently.
“With our proof of concept, we believe we’ve moved beyond the limitation of fixed power budgets,” said Dave Ward, president of cloud services at Lambda .