Source thumbnail for Baseten on the inference frontier: 200,000-token routing, 4–6× speedups, and self-optimizing AI
deep diveListen

Baseten on the inference frontier: 200,000-token routing, 4–6× speedups, and self-optimizing AI

This Latent Space podcast interview has host swyx asking Baseten’s Philip Kiely and Ali Taha how inference turns an open model into a product: request routing, quantization, GPU kernels, model parallelism, and AI video. It gets properly weird near the end, when GLM-5.2 helps rewrite its own serving code and continual learning starts to blur the line between training and inference.

  • AI Infrastructure, Compute, Chips, And Energy
  • AI Engineering, Software, And Developer Tooling
  • Open Models
The Inference Frontier: 10x Faster Models to Self-Optimizing AI — Philip Kiely & Ali Taha, Baseten
Source thumbnail for Jeff Dean on building useful AI systems
deep diveListen

Jeff Dean on building useful AI systems

Jeff Dean is one of the GOATs of modern computing: he helped Google scale, co-founded Google Brain, and is now Google’s Chief Scientist. At Y Combinator’s Startup School 2026, YC Managing Partner Diana Hu interviews him about agents working for weeks, the chips, memory, context and evaluation they need, and which assumptions and unsolved problems founders should chase.

  • AI Infrastructure, Compute, Chips, And Energy
  • Agents
  • AI Engineering, Software, And Developer Tooling
  • AI For Science
Jeff Dean: The 1% Rule for Building in AI
Source thumbnail for Alexandr Wang: “This is a Once-in-a-Civilization Opportunity”
deep diveListen

Alexandr Wang: “This is a Once-in-a-Civilization Opportunity”

This is a July 29, 2026 Y Combinator interview in which Garry Tan questions Scale AI founder and Meta Superintelligence Labs leader Alexandr Wang at Startup School 2026. It covers Scale's origin, Meta's frontier-lab rebuild, personal superintelligence, cheaper models, agent swarms, systems thinking, and Wang's advice to young builders. It matters because Wang gives a current frontier-lab operator's account of how AI may shift the constraint from access to intelligence toward the ability to direct large numbers of agents at useful goals.

  • Frontier Models And Capabilities
  • Agents
  • AI Engineering, Software, And Developer Tooling
  • Capital, Markets, And Business Models
Alexandr Wang: “This is a Once-in-a-Civilization Opportunity”
Source thumbnail for SemiAnalysis says strong AI demand can still run into leverage, cash-flow and financing limits
deep diveListen

SemiAnalysis says strong AI demand can still run into leverage, cash-flow and financing limits

This SemiAnalysis podcast episode is a discussion with Doug O'Loughlin about the July 2026 semiconductor stock drawdown and the AI infrastructure cycle. It covers leveraged memory trades, Chinese supply, token demand, AI politics, the timing gap between capital spending and revenue, debt markets, and local data-center constraints. It is important because strong AI usage can coexist with a financing or market correction if capital, power, labor, and political support do not scale as fast as infrastructure commitments. The episode's central distinction is between demand for AI and the path used to finance enough infrastructure to serve it. The speakers remain positive about model usage and memory demand. They also describe several ways the buildout can overshoot or stall before that demand produces enough cash.

  • AI Infrastructure, Compute, Chips, And Energy
  • Capital, Markets, And Business Models
SemiAnalysis
Source thumbnail for Core Automation says transformers cannot continually learn and is automating kernel generation to find a replacement
deep diveListen

Core Automation says transformers cannot continually learn and is automating kernel generation to find a replacement

This Recap covers a Sequoia Capital interview with Core Automation founders Jerry Tworek and Rohan Anil, hosted by Sonya Huang and Pat Grady. It is about the founders' claim that transformer architecture now limits further progress because deployed models cannot continually learn from real work. It matters because Core Automation is building an automated research lab to search for a replacement architecture, starting with the GPU kernels needed to run new ideas efficiently.

  • Frontier Models And Capabilities
  • AI Infrastructure, Compute, Chips, And Energy
  • AI For Science
Sequoia Capital
Source thumbnail for World Labs and SceniX are building digital worlds to train and evaluate robots
deep diveListen

World Labs and SceniX are building digital worlds to train and evaluate robots

This a16z interview brings together Martin Casado, World Labs co-founder and CEO Fei-Fei Li, and SceniX co-founder Yunzhu Li after World Labs acquired SceniX. It is about combining generative world models with a real-to-sim-to-real robotics pipeline so robots can learn and be evaluated inside controllable digital environments. It is important because robotics does not have the abundant internet-scale training data that accelerated language models, and reliable simulation could provide the scale, coverage, and measurement needed to move robot policies into real deployments. World Labs calls its field spatial intelligence: AI that can perceive, generate, reason about, and interact with physical or virtual spaces. SceniX approaches the same problem from the robotics side. Its aim is practical—build environments in which robot policies can train, fail, improve, and prove that they work before the corresponding hardware is trusted in the physical world.

  • World Models And Robotics
a16z
Source thumbnail for Sam Altman on AGI, Compute, and Human Agency
deep diveListen

Sam Altman on AGI, Compute, and Human Agency

This is a July 28, 2026 Invest Like The Best podcast interview in which Patrick O'Shaughnessy questions OpenAI CEO Sam Altman. It covers OpenAI's compute strategy, an AI-assisted cyber incident, Altman's AGI and jobs forecasts, personal agents, robotics, ChatGPT, and OpenAI's competitive position and structure. It matters because Altman connects the infrastructure race to the control people may retain as AI grows more capable, while naming security, governance, and economic problems OpenAI has not solved.

  • Frontier Models And Capabilities
  • AI Infrastructure, Compute, Chips, And Energy
  • Policy, Governance, And Geopolitics
  • World Models And Robotics
Sam Altman on AGI, Compute, and Human Agency
Thumbnail for Gavin Baker says GPU contract repricing can fund hyperscaler AI capex
deep dive

Gavin Baker says GPU contract repricing can fund hyperscaler AI capex

This is a July 28 X post by investor Gavin Baker arguing that the market is overreacting to wider credit spreads for hyperscalers such as Microsoft, Google, Amazon, and Meta. His thesis is that current cash flow understates what these companies will earn because much of their GPU capacity is still priced under older contracts. As those contracts expire and reset at higher rates, he expects operating cash flow to accelerate enough to fund the AI buildout, reduce the need for debt, and pull credit spreads back in.

  • AI Infrastructure, Compute, Chips, And Energy
  • Capital, Markets, And Business Models
Gavin Baker
Source thumbnail for Akshay Nathan says OpenAI is extending Codex from developers to knowledge work and personal agents
deep diveListen

Akshay Nathan says OpenAI is extending Codex from developers to knowledge work and personal agents

This Recap covers a Latent Space podcast interview with Akshay Nathan, who leads Core Product Engineering at OpenAI. It is about OpenAI's attempt to extend the agent architecture behind Codex into ChatGPT Work for knowledge work and personal productivity. It matters because OpenAI is combining persistent computers, plugins, artifacts, Sites, memory, scheduled tasks, and sub-agents into a general-purpose work interface—and Nathan says its success should be judged by completed goals, not AI-generated activity. OpenAI's current product guide describes ChatGPT Work as a place to delegate substantial tasks that produce reviewable outcomes. It says people who used Codex for non-coding work can stay in Codex or use the same core capabilities through an experience designed for everyday work. Nathan's interview explains why OpenAI made that choice and where it wants the product sequence to go next.

  • Agents
  • AI Engineering, Software, And Developer Tooling
  • Jobs, GDP, And Economic Growth
Latent Space
Source thumbnail for How contracts, utilization, software, and power shape AI infrastructure financing
deep diveListen

How contracts, utilization, software, and power shape AI infrastructure financing

This RAISE Summit 2026 panel is about who pays for AI infrastructure. IREN and Crusoe build data centres. Modular makes software that runs across different chips. Ornn is building a market for GPU capacity. Alva Energy upgrades existing nuclear plants to produce more power. AI infrastructure costs billions upfront. The panel explains what makes lenders say yes: customers who commit to buy, hardware that can be reused, and enough power to keep the machines running.

  • AI Infrastructure, Compute, Chips, And Energy
  • Capital, Markets, And Business Models
The Trillion Dollar Question: Financing AI's Infrastructure | Modular, IREN & More | RAISE 2026
Source thumbnail for Matt Murphy says open-source models will not displace Anthropic and AI outliers have reset venture investing
deep diveListen

Matt Murphy says open-source models will not displace Anthropic and AI outliers have reset venture investing

Matt Murphy came to Menlo Ventures in 2015 after 15 years at Kleiner Perkins. There, he was a Google board observer from the firm's initial investment through the IPO, launched the $200 million Apple-backed iFund, sourced DocuSign, and led investments including AppDynamics, which Cisco acquired for $3.7 billion, and Shazam, which Apple acquired. That history matters here: Murphy's career has repeatedly put him at the point where a new computing platform creates both an infrastructure layer and a new class of applications. Menlo's current $3 billion platform formalizes the approach: Menlo Ventures XVII invests from seed through Series A, while Menlo Inflection IV supplies growth capital from Series B onward. Anthropic became the anchor for the strategy. Murphy led Menlo's first investment in 2023, when Anthropic was pre-product and pre-revenue; Menlo later led its Series D with the largest check in the firm's history and says it has invested in every subsequent round. Menlo says its conviction came from the technical depth of Dario Amodei's team, room for another independent foundation-model company, and the view into infrastructure and applications created by proximity to the model layer. The other major investments discussed by Harry Stebbings fit that map. Menlo led OpenRouter's $40 million Series A because developers need one place to trade off model performance, price, latency, and provider capacity; co-led Lovable's $330 million Series B on the thesis that natural language expands who can create software; and joined Legora's $550 million Series D because high-context, trust-dependent legal workflows leave room for a purpose-built product above foundation models.

  • Frontier Models And Capabilities
  • Capital, Markets, And Business Models
  • Open Models
Will Open-Source Threaten Anthropic's Business & Do Margins Matter in a World of AI | Matt Murphy
Thumbnail for SSI says its research is ready to scale with NVIDIA
deep diveListen

SSI says its research is ready to scale with NVIDIA

Safe Superintelligence Inc. (SSI), the research lab Ilya Sutskever founded in 2024 to pursue safe superintelligence, announced a long-term partnership with NVIDIA on July 27, 2026. NVIDIA is investing in SSI and giving it access to Vera Rubin systems. SSI says this will increase its compute tenfold over 12 months. The lab says its research is ready to scale but has not disclosed the research.

  • Frontier Models And Capabilities
  • AI Infrastructure, Compute, Chips, And Energy
  • Capital, Markets, And Business Models
SSI and NVIDIA announce long-term strategic partnership
Thumbnail for Nvidia and OpenAI discuss a $250 billion infrastructure backstop
deep dive

Nvidia and OpenAI discuss a $250 billion infrastructure backstop

CNBC reporters Ashley Capoot and Kate Rooney describe Nvidia and OpenAI's talks over a proposed financing backstop for a 10-gigawatt data-center campus in Ohio. The article covers the debt Nvidia may support, the scale of the campus, and the companies already financing OpenAI's infrastructure.

  • AI Infrastructure, Compute, Chips, And Energy
  • Capital, Markets, And Business Models
Nvidia and OpenAI in talks for up to $250 billion backstop to fund AI infrastructure plans
OPEN-WEIGHTS BANNED? — Brad beside the headline, looking directly at the camera.
publication

The open weights debate (3D Chess)

Since the launch of GPT-3, AI has become an increasingly complicated game of geopolitical and commercial 3D chess. Everyone can see the immense potential benefits of the technology. Everyone can also see the risks that come with falling behind. But the participants in this race do not all want the same thing. The labs want better models and a lead in the race to AGI. Governments want security, power and stability. Investors want returns. Hardware and infrastructure providers want more demand and wider adoption of their technology.

  • Open Models
  • Policy, Governance, And Geopolitics
  • Capital, Markets, And Business Models
Source thumbnail for NVIDIA’s Central Bank of AI
deep diveListen

NVIDIA’s Central Bank of AI

This is a SemiAnalysis Weekly panel with Jordan Nanos, Dan Nishball, Zane Fong, and Kang Wen Cheang about financing AI infrastructure. It covers the capital–offtake–data-center trinity, NVIDIA’s GPU debt backstops, neocloud lending, and the token economics beneath the buildout.

  • AI Infrastructure, Compute, Chips, And Energy
  • Capital, Markets, And Business Models
SemiAnalysis
Source thumbnail for Sam Altman — How to Start a Startup
deep dive

Sam Altman — How to Start a Startup

This is Ti Morse’s July 2026 Relentless interview with OpenAI CEO Sam Altman about building startups and operating OpenAI. They discuss AI-era company formation, founder psychology, execution, compute, OpenAI’s evolution, product design, leadership, and durable shifts in user behavior.

  • AI Infrastructure, Compute, Chips, And Energy
  • Agents
  • Capital, Markets, And Business Models
Sam Altman - How to Start a Startup
Source thumbnail for Kimi K3 Is the Best Model Ever Made — Sometimes
deep dive

Kimi K3 Is the Best Model Ever Made — Sometimes

This is Theo Browne’s hands-on video review of Moonshot AI’s Kimi K3 after a day spent pushing the model through coding, frontend, 3D, research, security, and multi-agent work. It covers why K3 feels like a frontier-class open-weight release, the architecture and economics behind it, the work it did well, and the rough interaction layer behind the “sometimes” in Theo’s title.

  • Frontier Models And Capabilities
  • AI Infrastructure, Compute, Chips, And Energy
  • Agents
  • Open Models
Kimi K3 is the best model ever made (sometimes)
Source thumbnail for Poolside’s Model Factory, Laguna S, Open Models, and the Race to AGI — Eiso Kant, Poolside AI
deep diveListen

Poolside’s Model Factory, Laguna S, Open Models, and the Race to AGI — Eiso Kant, Poolside AI

This Recap covers Latent Space’s YouTube interview with Eiso Kant, co-founder of Poolside, an AI lab building models for long, complex coding work; a Baseten interview and Poolside’s own releases add supporting detail. The company recently launched Laguna S 2.1, a coding model released with open weights and designed to reason through difficult tasks over many steps. In the interview, Kant explains the Model Factory—the data, training, evaluation, and engineering system behind Poolside’s models—and Poolside’s philosophy of model building, shaped by years of doing the work. His argument is that open research should share how the factory works and why it was built that way, not just the finished weights and benchmarks.

  • Frontier Models And Capabilities
  • AI Infrastructure, Compute, Chips, And Energy
  • Agents
  • Open Models
Latent SpaceBasetenLaguna XS.2 and M.1: A Deeper Dive
Source thumbnail for Poolside Co-Founder Explains Frontier Open-Weight Models
deep dive

Poolside Co-Founder Explains Frontier Open-Weight Models

This 34-minute Baseten interview with Poolside co-founder Eiso Kant explains Laguna S 2.1 and Poolside's push to keep frontier model weights open. Gavin Baker's response, Jason Warner's launch post, and Poolside's release article cover the architecture, benchmarks, training system, public trajectories, and deployment.

  • Frontier Models And Capabilities
  • AI Infrastructure, Compute, Chips, And Energy
  • Agents
  • Open Models
Baseten
Thumbnail for Who’s Afraid of Chinese Models?
deep dive

Who’s Afraid of Chinese Models?

This is Ben Thompson’s July 20, 2026 Stratechery analysis of Chinese open-weight AI models. It covers inference economics, frontier-lab strategy, China’s industrial policy, distillation, and cybersecurity.

  • Frontier Models And Capabilities
  • Capital, Markets, And Business Models
  • Policy, Governance, And Geopolitics
  • Open Models
Who’s Afraid of Chinese Models?
Thumbnail for Kimi K3: The open-weights escalation
deep dive

Kimi K3: The open-weights escalation

Nathan Lambert argues that Kimi K3 makes frontier open weights credible, pressuring closed-lab economics while forcing a harder global policy and risk-management problem.

  • Frontier Models And Capabilities
  • Capital, Markets, And Business Models
  • Policy, Governance, And Geopolitics
  • Open Models
Nathan Lambert
Source thumbnail for Conversation With Baseten's Head of Training
deep dive

Conversation With Baseten's Head of Training

Charles O'Neill argues that teams should prove a task with a strong hosted model, then use post-training and open weights when proprietary feedback, control, cost, or latency justify specialization.

  • AI Infrastructure, Compute, Chips, And Energy
  • Open Models
Baseten
Thumbnail for From AGI to ASI
deep dive

From AGI to ASI

A Google DeepMind-led report maps four possible routes from AGI to ASI—scaling, new algorithms, recursive improvement, and AI collectives—and the frictions that could stall each one.

  • Frontier Models And Capabilities
Google DeepMind
Thumbnail for AI 2040: Plan A
deep dive

AI 2040: Plan A

AI Futures Project proposes a verified US-China deal that opens frontier AI research, controls compute, pauses at top-expert AI, and delays superintelligence until 2040.

  • Frontier Models And Capabilities
  • AI Infrastructure, Compute, Chips, And Energy
  • Jobs, GDP, And Economic Growth
  • Policy, Governance, And Geopolitics
AI Futures Project
Source thumbnail for AI Sputnik Moment: Kimi K3
deep dive

AI Sputnik Moment: Kimi K3

Kimi K3 reached the frontier on price and capability, while several claims about its rank, open weights, training economics, talent flows, forecasting, and screening outran the available evidence.

Urgent Update- AI Sputnik Moment: Kimi K3 Released w/ Emad Mostaque | Ep. 272
Ilya Sutskever and Demis Hassabis with the words Continual Learning
publication

Why We May Never ‘Solve’ Continual Learning

Continual learning, It’s a defining but poorly understood feature of biological intelligence. In artificial systems, we do not yet know what it would take to reproduce this capability, if it's even possible to replicate in a biologically-similar way, or how essential it is to further AI progress. In the end, it might turn out that there was one clear path to solving continual learning. Maybe that path will lead us to AGI and beyond. And maybe we’re on that path now. It's also possible that there will be many paths to progress.

  • Frontier Models And Capabilities
  • Agents
Source thumbnail for Bryan Catanzaro On Why NVIDIA Builds Nemotron
deep dive

Bryan Catanzaro On Why NVIDIA Builds Nemotron

  • Frontier Models And Capabilities
  • AI Infrastructure, Compute, Chips, And Energy
  • Agents
  • AI Engineering, Software, And Developer Tooling
  • Jobs, GDP, And Economic Growth
  • Capital, Markets, And Business Models
  • Policy, Governance, And Geopolitics
  • Open Models
The MAD Podcast
Thumbnail for Lilian Weng Explains Why Scaling Laws Need Careful Accounting
deep dive

Lilian Weng Explains Why Scaling Laws Need Careful Accounting

Weng explains how scaling laws help labs plan large AI training runs, then shows why Kaplan and Chinchilla gave different compute-allocation advice.

  • Capital, Markets, And Business Models
  • AI Engineering, Software, And Developer Tooling
  • Frontier Models And Capabilities
  • AI Infrastructure, Compute, Chips, And Energy
Lilian Weng
Source thumbnail for Stephen Balaban on the energy-to-tokens machine behind AI compute
deep dive

Stephen Balaban on the energy-to-tokens machine behind AI compute

Stephen Balaban starts the story before the GPU. In Matt Turck's interview, Lambda's co-founder and CTO says the clean way to think about AI compute is to put energy production on the left and tokens on the right. Photons from the sun, or molecules of natural gas, become electrical power. The data center consumes those watts, spends some of them cooling itself, and feeds the rest into servers, networking, and storage. That hardware produces FLOPS. Model builders consume those FLOPS for training and inference. Then the FLOPS become tokens per second, and the application tries to turn those tokens into useful intelligence. That is why a customer does not log into an AI cloud and buy a loose GPU off a shelf. After the energy reaches the rack, Balaban says the hard part is turning the chips into a real cloud service. He asks listeners to imagine a 10,000-GPU cluster. Then the work begins. The cluster needs storage fast enough for model data. It needs CPU servers to orchestrate jobs. It needs one network for ordinary traffic, another for monitoring, and a compute fabric where GPUs can move model weights and activations between machines. When Lambda partitions that system for a customer, all of those layers have to move together. Balaban calls it an immense software undertaking, and says many neoclouds have not made the investment needed to run a real cloud service. That is the GPU myth. The product is not the chip. The product is the machine that makes the chip usable.

  • AI Infrastructure, Compute, Chips, And Energy
  • Capital, Markets, And Business Models
The MAD Podcast
The AGI Post
publication

AI’s New Cost Curve

AI is shifting scarcity from raw capability to the costs around it: inference budgets, chip data movement, reliable evaluation, durable memory, human supervision, capital substitution for labor, and weak-link institutions. This issue maps w

Latent.SpaceOpenAIswyx