Stephen Balaban on the energy-to-tokens machine behind AI compute
- Capital, Markets, And Business Models
- AI Infrastructure, Compute, Chips, And Energy
Watch the deep dive
The big thing is that cloud compute is not a commodity service. It is a very complicated highly vertically integrated type of service.
Matt Turck interviewed Stephen Balaban, Lambda co-founder and CTO, in a June 18, 2026 episode of The MAD Podcast about why AI compute did not become simple GPU rental. Balaban's core claim is that an AI cloud has to join entitled land, construction, power, cooling, high-performance computing design, networking, storage, virtualization, customer-facing cloud software, financing, and customer demand into one usable service. That is why he says GPU rental indexes can mislead when they mix short-term on-demand prices with long-term contracts: the harder question is whether a provider can turn chips into reliable, partitioned, customer-ready compute.
Demand is the reason Balaban thinks the market is still generally underbuilt. If better models make more work worth doing with AI, compute demand can expand with capability. Even a 10x efficiency gain may simply let customers process 10x more tokens with the same fixed compute, where tokens are the units models process and generate. The bottleneck then becomes physical as well as technical: land, power, data-center shell, generators, UPS systems, and mechanical, electrical, and plumbing gear. Balaban says communities have real concerns about large projects, while arguing that some data-center criticism overstates water use because modern direct-to-chip liquid cooling can use dry coolers with very low evaporation.
His clearest explanation is the energy-to-tokens chain. Energy inputs become electrical power; the data center loses some power to cooling and overhead, often tracked through PUE, or Power Usage Effectiveness; servers, networking, and storage turn the remaining power into FLOPS, the raw operations used for training and inference; and model builders then care how much of that theoretical compute does useful work, a measure Balaban discusses through MFU, or Model FLOPs Utilization. At the end, FLOPS become tokens per second, and the application tries to turn those tokens into useful intelligence.
The platform layer is what makes the chip usable. Balaban says NVIDIA's moat is not only silicon but CUDA, cuDNN, networking, system reliability, and developer ecosystem gravity. Lambda's one-click cluster product is the same idea from the cloud provider side: the customer sees a simpler abstraction, while the provider has to coordinate GPU servers, CPU orchestration servers, storage, ordinary traffic networks, monitoring networks, and a compute fabric that can move data between GPUs without unnecessary CPU copies. Balaban calls that an immense software undertaking.
The finance section turns the same thesis into a balance-sheet story. Balaban says customer off-take agreements, special-purpose vehicles, private credit, and asset-backed lending can finance GPU deployments when lenders trust the customer cash flows and the NVIDIA chip assets. His sharpest example is that some H100s deployed in 2023 can lease for more later because demand stayed high and useful life looked longer than skeptics expected. He then connects that infrastructure business back to Lambda's origin story, from facial-recognition work and hardware experiments to a roughly $60,000 GPU purchase that became a cloud wedge, and forward to a future of neural software, agents, gigawatt-scale AI factories, and compute broad enough to support "one person, one GPU."
Section 1
Section 01- Matt Turck interviewed Stephen Balaban, Lambda co-founder and CTO, in a June 18, 2026 episode of The MAD Podcast about why AI compute did not become simple GPU rental. Balaban's core claim is that an AI cloud has to join entitled land, construction, power, cooling, high-performance computing design, networking, storage, virtualization, customer-facing cloud software, financing, and customer demand into one usable service. That is why he says GPU rental indexes can mislead when they mix short-term on-demand prices with long-term contracts: the harder question is whether a provider can turn chips into reliable, partitioned, customer-ready compute.
Section 2
Section 02- Demand is the reason Balaban thinks the market is still generally underbuilt. If better models make more work worth doing with AI, compute demand can expand with capability. Even a 10x efficiency gain may simply let customers process 10x more tokens with the same fixed compute, where tokens are the units models process and generate. The bottleneck then becomes physical as well as technical: land, power, data-center shell, generators, UPS systems, and mechanical, electrical, and plumbing gear. Balaban says communities have real concerns about large projects, while arguing that some data-center criticism overstates water use because modern direct-to-chip liquid cooling can use dry coolers with very low evaporation.
Section 3
Section 03- His clearest explanation is the energy-to-tokens chain. Energy inputs become electrical power; the data center loses some power to cooling and overhead, often tracked through PUE, or Power Usage Effectiveness; servers, networking, and storage turn the remaining power into FLOPS, the raw operations used for training and inference; and model builders then care how much of that theoretical compute does useful work, a measure Balaban discusses through MFU, or Model FLOPs Utilization. At the end, FLOPS become tokens per second, and the application tries to turn those tokens into useful intelligence.
Section 4
Section 04- The platform layer is what makes the chip usable. Balaban says NVIDIA's moat is not only silicon but CUDA, cuDNN, networking, system reliability, and developer ecosystem gravity. Lambda's one-click cluster product is the same idea from the cloud provider side: the customer sees a simpler abstraction, while the provider has to coordinate GPU servers, CPU orchestration servers, storage, ordinary traffic networks, monitoring networks, and a compute fabric that can move data between GPUs without unnecessary CPU copies. Balaban calls that an immense software undertaking.
Section 5
Section 05- The finance section turns the same thesis into a balance-sheet story. Balaban says customer off-take agreements, special-purpose vehicles, private credit, and asset-backed lending can finance GPU deployments when lenders trust the customer cash flows and the NVIDIA chip assets. His sharpest example is that some H100s deployed in 2023 can lease for more later because demand stayed high and useful life looked longer than skeptics expected. He then connects that infrastructure business back to Lambda's origin story, from facial-recognition work and hardware experiments to a roughly $60,000 GPU purchase that became a cloud wedge, and forward to a future of neural software, agents, gigawatt-scale AI factories, and compute broad enough to support "one person, one GPU."