Dwarkesh Patel on How AI Agents Could Learn on the Job
- Jobs, GDP, And Economic Growth
- World Models And Robotics
- Agents
- AI Engineering, Software, And Developer Tooling
- Frontier Models And Capabilities
Watch the deep dive
If AIs are to develop all the skills that humans have, and even skills that humans don't have, then they need to be able to learn from information revealed in unstructured, unverifiable, and ambiguous ways from scarce amounts of real-world interaction. Because in many domains, the relevant training information simply doesn't exist in any other way.
Dwarkesh Patel says frontier labs are betting on RLVR: train agents on millions of tasks with checkable answers until they become broad problem solvers. His doubt is that verifiable is not enough. A task also has to be grindable, meaning it can be replayed many times from the same starting point. Coding can work like that. Business, politics, law, markets, and operations usually cannot.
Section 1
Section 01- Dwarkesh Patel says frontier labs are betting on RLVR: train agents on millions of tasks with checkable answers until they become broad problem solvers. His doubt is that verifiable is not enough. A task also has to be grindable, meaning it can be replayed many times from the same starting point. Coding can work like that. Business, politics, law, markets, and operations usually cannot.

Tags
- Coding Agents
- Labor Automation
- Frontier Models
- Test-Time Compute
- Software Reliability And Verification
- Workflow Automation
- Benchmarks And Evaluation
- Human-In-The-Loop Agents
- Post-Training