Discussion about this post

User's avatar
Nitin's avatar

thank you for another very helpful distillation of very confusing data points.

and i think your article is a perfect example of how a human, with clear thinking and good articulation, can balance so many conflicting ideas, their biases incentives blind spots etc, and collate them into something that moves the ball forward while keeping the possibilities open. despite LLMs being great at absorbing vastly more information and great at connecting dots, and frequently doing a great job at a similar task, i 'feel' something is missing in their analysis, or more accurately, being more than required confident and therefore preemptively closing some paths in a subtle way. and i think one factor that leads to this difference is humans knack for knowing what they don't know. there are other factors too like humans know how they have been wrong, in small and big ways, in their analysis and predictions so many times in their life, and that brings in a natural self-doubt/humility. an LLM has some notion of to be careful and not being overly confident but it hasn't grown this self-doubt in a natural way and is unsurprisingly artificial and not as effective.

all of this is i guess to say that any research/exploration requires at least 3 things - ability to connect dots, self-doubt, and confidence to own the outcome of decision. LLMs lack the last and self-doubt is not very effective. massive parallel search can compensate for some of the limitations but not sufficiently, especially the lack of ownership/accountability.

massive parallel search approach can hide this limitation, especially when there is a verifiable right answer but because it is almost impossible, even with current compute, to study and learn from the path taken to get there, all failed/dead-end searches do not help with any other pursuit. this is very different from my understanding of how science progresses - people try a much smaller number of paths but learn from many of them, and ideas from one field (even of failed experiments) trigger ideas in other related and unrelated fields. so even though on one particular problem the progress might seem slow (compared to AI), the process is much more organic/natural and therefore more robust to sustain and expand growth in many directions.

for these reasons it seems to me that AI is like a laser that can really lit up one very tiny surface very brightly, but when exploring what is more efficient and effective over long time horizon is a much dimmer torch that sheds diffused light over a much larger area.

Sam's avatar

I appreciate your measured take, and I agree with most of it.

> Diminishing returns of more AI agents in parallel are real

I can't help but feel that we are only in the foothills of scaling agents. I agree that Amdahl's Law is likely in play here -- I think Noam said they don't yet have good heuristics on what number/structure of agents works best. My guess is there is still a lot of capability left to unlock as they iterate on this paradigm. Right now we see the progress in the verifiable domains, but I bet we will see some generalization in other less verifiable domains. When and how much, I'm not sure.

Mainly, I'm blown away that they had 10k agents work towards a problem *at all* without it degenerating into a hallucination spiral. To me that's the biggest surprise out of all of this

7 more comments...

No posts

Ready for more?