Writing

Reward models are quietly becoming the most important part of LLM progress

Pre-training is slowing down as the internet stops growing as fast as new LLMs launch. Maybe today's new releases are less new base models, more better reward models.

Reward models are quietly becoming the most important part of LLM progress.

Pre-training is starting to slow down, the internet isn’t growing as fast as big tech companies are launching new LLMs. Pre-training mainly helps models understand how language works, but not how humans think or prefer responses.

The better the reward model, the smarter the LLM.

Maybe the new version releases of LLMs we see today aren’t always new base models, just better human+benchmark aligned versions built with stronger reward models.