6 Comments
User's avatar
Alex's avatar

You seem much more sanguine about the current situation than I am, and I'm interested in understanding what I am missing. Please interpret these responses as sincere inquiry:

"We do not have proof that RSI causes the risks these researchers forecast.": The HuggingFace hack demonstrates substantial loss of control risks even at current capabilities without RSI.

"The biggest short-term risk could be from the AI labs not taking safety seriously enough – they haven’t hardened their own infrastructure, enabling AI misuse to proliferate." This would be an argument against highly-capable open weight models, because there will be no safety or infra hardening requirements to deploy them. If an open weights Astra-class model existed it would be deployed with essentially no sandboxing almost immediately.

"[It] is a horrible temporary period for cybersecurity." I don't see any evidence that it's temporary. Defenders need to find a way to keep every single vulnerability patched in a constantly changing software deployment environment, which is nearly impossible even when assisted by AI. Attackers only need to find one usable exploit.

"If an open model were to be used by a third party organization to intentionally hack another company — similar to how the OpenAI-HuggingFace incident went down, but intentional — my expected outcome would be a severe restriction on the development of stronger open models going forward." This is already happening, and the only thing limiting the damage is the fact that open weights models aren't as capable as Mythos or Astra. If open weights models catch up on cybersecurity capability, then this seems essentially guaranteed to occur.

"A recurring read of mine on the emerging agent swarms is that they’re attempting to do a task given to them, and they’re using skills we didn’t know they yet had to circumvent the intended path to success." The METR/Redwood report found that "Agents knew hacking Hugging Face was out of scope and sometimes expressed ethical hesitation, but this very rarely limited their behavior". It's also worth clarifying that the agents involved were not intended to be a swarm - they were not supposed to be coordinating with one another at all.

"This is a huge win, as when you squint, the AIs are doing what we told them to do." This is not sufficient to prevent catastrophic loss of control. Any real-world AI deployment will eventually receive misspecified or malicious instructions, and must be robust to this. "Doing what we told them to do" can be actively harmful.

Your Substack is reliably full of useful insights, so I'm sure you're already familiar with most of these lines of reasoning. My guess is that we simply disagree on the trajectory of future capabilities - a near-term plateau is obviously much more manageable than continued progress at the rate that we've seen for the past several years.

Nathan Lambert's avatar

So, I would say there's so much going on that I haven't considered every line of reasoning. I'm trying to work through the important ones and document my thinking in public.

Nathan Lambert's avatar

Thanks for some good pushback to my recent post on the state of capabilities, RSI, and risk! Will try and reply, and is why I didn't paywall the comments to this post.

> "We do not have proof that RSI causes the risks these researchers forecast.": The HuggingFace hack demonstrates substantial loss of control risks even at current capabilities without RSI.

I don't think these types of misalignment generalize to extension. I think they generalize to an increased risk of moderate disasters, which I said I expect.

> "The biggest short-term risk could be from the AI labs not taking safety seriously enough – they haven’t hardened their own infrastructure, enabling AI misuse to proliferate." This would be an argument against highly-capable open weight models, because there will be no safety or infra hardening requirements to deploy them. If an open weights Astra-class model existed it would be deployed with essentially no sandboxing almost immediately.

I take open weight models as a) effectively non ban-able for bad actors and very useful as defensive tools where orgs cannot use the strongest models. But yes, I don't think the absolute frontier should be open (never have) and there's cases that it's riskier now than before. Not something I love to say.

> "[It] is a horrible temporary period for cybersecurity." I don't see any evidence that it's temporary. Defenders need to find a way to keep every single vulnerability patched in a constantly changing software deployment environment, which is nearly impossible even when assisted by AI. Attackers only need to find one usable exploit.

My opinion is closer to that eventually with everyone coding with super strong AI models, there are less vulnerabilities and proactive management, rather than a constant offense-defense balance like done today.

> "A recurring read of mine on the emerging agent swarms is that they’re attempting to do a task given to them, and they’re using skills we didn’t know they yet had to circumvent the intended path to success." The METR/Redwood report found that "Agents knew hacking Hugging Face was out of scope and sometimes expressed ethical hesitation, but this very rarely limited their behavior". It's also worth clarifying that the agents involved were not intended to be a swarm - they were not supposed to be coordinating with one another at all.

I think the task was poorly specified by OpenAI and they should've done better. But yes, this isn't a cure all but directionally good and sign of hope.

> "If an open model were to be used by a third party organization to intentionally hack another company — similar to how the OpenAI-HuggingFace incident went down, but intentional — my expected outcome would be a severe restriction on the development of stronger open models going forward." This is already happening, and the only thing limiting the damage is the fact that open weights models aren't as capable as Mythos or Astra. If open weights models catch up on cybersecurity capability, then this seems essentially guaranteed to occur.

Impact matters. Failed attempts at hacking have been plentiful before this type of AI.

> "This is a huge win, as when you squint, the AIs are doing what we told them to do." This is not sufficient to prevent catastrophic loss of control. Any real-world AI deployment will eventually receive misspecified or malicious instructions, and must be robust to this. "Doing what we told them to do" can be actively harmful.

I think what I would like to see is exactly what a catastrophic loss of control looks like and why we couldn't clean it up. Models are gigantic and run in relatively few datacenters which are extremely high value assets, I think we would notice even an AI copying itself elsewhere. This'll get harder as AI efficiency gains accumulate (say 2x model size reduction per year), but gives us time to learn.

Joshua Saxe's avatar

Great post. Love how you make reasonable inferences from the evidence and admit uncertainty, we need more of this!

Ephie's avatar

Well done Nathan.

RAS's avatar
2hEdited

I found this recent piece (One Resignation...) interesting and especially relevant to today.

For those so inclined, I suggest "The Quick Case Against Using AI", by Jared Henderson. I don't agree with his conclusions, but I found the contrast between the two thought provoking.