Discussion about this post

User's avatar
Skillselion's avatar

The grader framing also predicts the ordering of the next dominoes, and the axis that matters is reward latency, more than verifiability itself. Code returns a verdict in seconds. Search rankings return one in weeks, sales in months, strategy in years: the signal exists but arrives too slowly and too confounded to train against cheaply. So the domains that fall next are the ones where someone compresses the loop - simulators, proxy metrics that track the real outcome closely enough, or environments that replay historical ground truth as if live. Which suggests "who writes the next answer key" is really "who can make a slow answer key fast", and that is a different, harder engineering problem than grading code.

Rakesh Cheerla's avatar

Since the value is in the harness and post-training, this "whoever drives the grader wins" capability diffuses relatively quickly into the AI ecosystem? Does that mean that coding becomes (or has become) a commodity, not tied to the capabilities of a frontier closed model?

3 more comments...

No posts

Ready for more?