thank you so much for this!! really fortuitous timing bc I was just trying to collect my own reading list on open-weight models, but this is a much better starting point—really appreciate you sharing
To what extent is the 4 month gap simply size? If you step away from frontier sized training and model, it looks like open weights models win at every size including past proprietary frontier models. Are they actually leading in some of the technology for handling long contexts or aggressive quantization, in the sense that if they wanted to slow the rate of iteration they could train even larger and be more of a challenge to the frontier? Some of what the iterations are delivering are faster, cheaper, more stable methods of pretraining.
Can approve, very helpful list, I had a few of them already bookmarked!
thank you so much for this!! really fortuitous timing bc I was just trying to collect my own reading list on open-weight models, but this is a much better starting point—really appreciate you sharing
To what extent is the 4 month gap simply size? If you step away from frontier sized training and model, it looks like open weights models win at every size including past proprietary frontier models. Are they actually leading in some of the technology for handling long contexts or aggressive quantization, in the sense that if they wanted to slow the rate of iteration they could train even larger and be more of a challenge to the frontier? Some of what the iterations are delivering are faster, cheaper, more stable methods of pretraining.
lovely, time for students to read this instead (or alongside!) of the AI safety readings. Thank you very much
Awesome .... on behalf of lay people!
I wish I had time to read all of these!