Discussion about this post

User's avatar
Basil Wong's avatar

I found this article especially useful because you wrote it from the frame of: "comparing what the paper talks about to how pretraining scaling laws are used"!

Since you posted this article there have been a few papers released on the scaling laws of RL. Would be super interested if you ever had the opportunity to post a follow up article on the evolution of RL scaling. Especially with the growing usage of MOPD for large models...

Ram  Komarraju's avatar

Nathan, IIRC, you're quite bullish about the long term prospects of RL as applied to LLMs. But what do you think of the findings from the paper published last week "Reasoning with Sampling: Your Base Model is Smarter Than You Think" which seems to further confirm the results of the pass@k paper from earlier in the year. Is the only advantage of RLVR is improved one-shot performance in exchange for the loss of diversity? And even this is only applicable to verifiable scenarios?

Thanks

2 more comments...

No posts

Ready for more?