Really enjoyed this. The verification evolving from simple benchmarks to reward models to hybrid human-AI loops to "living systems that resist reward hacking" is basically a compressed replay of reinforcement learning's own sixty-year history. RL went through the exact same arc: hardcoded rewards (Atari scores), then learned reward models (RLHF), then the realization you need something like a world model to predict consequences and self-verify. And each stage broke the same way. Goodhart's Law. Optimize hard against any fixed signal, the agent finds the exploit. The line about "graders drift, benchmarks saturate, reward hacking appears" is literally the specification gaming problem RL researchers have been wrestling with since the 1990s, just showing up in a new costume.
Which pushes toward an interesting endpoint. To verify whether an action is "correct," you need to predict its consequences. That's a world model. The RL environment companies that win are the ones that end up building the best world models of their workflows, whether they call it that or not. The market map here might actually be a subset of a much bigger one.
Thanks for writing this, it clarifies a lot. The semiconductor analogy is incredibly insightful. It really makes me reflect on how this 'verification problem' keeps reappearing. This feels like a natural next step after your piece on AI scaling; it’s teh crucial maturity layer we need for agentic AI.
Really enjoyed the piece. One question I kept coming back to is the "lab first → enterprise later" thesis.
Is there a bit of a chicken-and-egg problem here?
If enterprises simply consume RL infrastructure (envs in particular) developed for frontier labs which will produce that resulting in significant economic uplift, then the labs need enterprise workflows to learn from in the first place. But many of the most valuable long-horizon workflows are deeply company-specific and aren't available to the labs in the first place.
That makes me wonder whether the real bottleneck isn't just replication training or verification (still important absolutely), but the process of extracting the latent semantics from real enterprise work and turning them into reproducible RL environments.
If that capability exists, it seems plausible that enterprises themselves could become continuous producers of RL environments, rather than only downstream consumers of lab-built infrastructure.
Curious how you think about that dynamic. Do you expect most of the long-term learning loop to remain centralised in the labs, or do you think enterprises will eventually own meaningful parts of the training infrastructure themselves? Or giving birth to meaningful training tasks
"But many of the most valuable long-horizon workflows are deeply company-specific and aren't available to the labs in the first place."
I absolutely agree with this point as key. Some companies are buying data sets from failed startups, others are trying novel ways to capture workflow data.
I think most of the short to medium term value is in the labs and that gives rise to a number of large players, who then have that enterprise opportunity as the second act -- and the second act trumps the first act long term.
Workflow capturing is super important here. Buying the failed startups only get you to certain point and from there you failed to capture evolving new knowledge that humans come up with. Though some may argue that for many human work, it's repetitive anyways. Anyways, quite bullish on novel capturing mechanisms and how AI-human coordination play out there inside the capturing workflow.
Really enjoyed this. The verification evolving from simple benchmarks to reward models to hybrid human-AI loops to "living systems that resist reward hacking" is basically a compressed replay of reinforcement learning's own sixty-year history. RL went through the exact same arc: hardcoded rewards (Atari scores), then learned reward models (RLHF), then the realization you need something like a world model to predict consequences and self-verify. And each stage broke the same way. Goodhart's Law. Optimize hard against any fixed signal, the agent finds the exploit. The line about "graders drift, benchmarks saturate, reward hacking appears" is literally the specification gaming problem RL researchers have been wrestling with since the 1990s, just showing up in a new costume.
Which pushes toward an interesting endpoint. To verify whether an action is "correct," you need to predict its consequences. That's a world model. The RL environment companies that win are the ones that end up building the best world models of their workflows, whether they call it that or not. The market map here might actually be a subset of a much bigger one.
Insightful comment! Thank you.
Great read thank you for sharing your insights!
Thanks for writing this, it clarifies a lot. The semiconductor analogy is incredibly insightful. It really makes me reflect on how this 'verification problem' keeps reappearing. This feels like a natural next step after your piece on AI scaling; it’s teh crucial maturity layer we need for agentic AI.
well said! completely agree.
Great writeup! Loved it
Really enjoyed the piece. One question I kept coming back to is the "lab first → enterprise later" thesis.
Is there a bit of a chicken-and-egg problem here?
If enterprises simply consume RL infrastructure (envs in particular) developed for frontier labs which will produce that resulting in significant economic uplift, then the labs need enterprise workflows to learn from in the first place. But many of the most valuable long-horizon workflows are deeply company-specific and aren't available to the labs in the first place.
That makes me wonder whether the real bottleneck isn't just replication training or verification (still important absolutely), but the process of extracting the latent semantics from real enterprise work and turning them into reproducible RL environments.
If that capability exists, it seems plausible that enterprises themselves could become continuous producers of RL environments, rather than only downstream consumers of lab-built infrastructure.
Curious how you think about that dynamic. Do you expect most of the long-term learning loop to remain centralised in the labs, or do you think enterprises will eventually own meaningful parts of the training infrastructure themselves? Or giving birth to meaningful training tasks
"But many of the most valuable long-horizon workflows are deeply company-specific and aren't available to the labs in the first place."
I absolutely agree with this point as key. Some companies are buying data sets from failed startups, others are trying novel ways to capture workflow data.
I think most of the short to medium term value is in the labs and that gives rise to a number of large players, who then have that enterprise opportunity as the second act -- and the second act trumps the first act long term.
Workflow capturing is super important here. Buying the failed startups only get you to certain point and from there you failed to capture evolving new knowledge that humans come up with. Though some may argue that for many human work, it's repetitive anyways. Anyways, quite bullish on novel capturing mechanisms and how AI-human coordination play out there inside the capturing workflow.