@sleepinyourhat - More in my new essay on the Anthropic alignment blog here
Sam Bowman@sleepinyourhat
We'll need to do a very good job at aligning the early AGI systems that will go on to automate much of AI R&D.
Our understanding of alignment is pretty limited, and when the time comes, I don't think we'll be confident we know what we're doing.
image not captured
Sam Bowman@sleepinyourhat
This isn't a good situation, but I think we're in a much better shape than this makes it sound. I don't think we need a breakthrough: It seems likely that there are at least *some* ways of using current techniques that would do the job.
Read 10 more tweets
Sam Bowman@sleepinyourhat
We haven't failed badly yet. Current systems are reasonably well-aligned. When we do see misalignment, it's usually pretty clear what caused it.
We’re not qualitatively that far from human-level R&D skills, and our situation when we get there might look a lot like this.
Sam Bowman@sleepinyourhat
More importantly, I think catching misalignment is tractable, at least for models up to around human level. We have a lot of tools at our disposal—red-teaming, monitoring, honeypots, interpretability, fuzzing, counterfactual prompts.
Sam Bowman@sleepinyourhat
I don’t expect these tools to all fail together if we use them well. Even if models try to actively evade our attempts to test them for misalignment, we have a lot of advantages that I expect will allow us to succeed anyway with a very thorough effort.
Sam Bowman@sleepinyourhat
I propose that we make that very thorough effort.
My proposal: Accept early AGI alignment will involve a good deal of trial and error, but aim to be excellent at spotting the errors.
Sam Bowman@sleepinyourhat
We’re not there yet. This will take practice, engineering, and at least a bit of new science. But I don’t think it’ll take any breakthroughs.
Sam Bowman@sleepinyourhat
The basic logic is simple, and a bit scary, but I think it’s workable:
1) Try to align the model as best we can with existing methods
2) Test it for misalignment
3) If we find something, diagnose and adjust—try the first intervention that seems to get at the root of the problem
Sam Bowman@sleepinyourhat
4) Return to (2)
5) Repeat until we no longer find anything
6) Put monitoring and control measures in place after deployment, if they find anything, return to (2)
Sam Bowman@sleepinyourhat
This isn't costless—each iteration creates some selection pressure against our tests. And it gives up on highly confident safety arguments. But in practice, it’s a good bet.
Sam Bowman@sleepinyourhat
I think it fails gracefully if it fails—more so than trying to deeply solve the problem in time, and more so than focusing exclusively on trying to control systems that we know are misaligned.
Sam Bowman@sleepinyourhat
This isn’t the only bet Anthropic is making for how to handle AGI safety, but it’s a big one. We're hiring across several teams—two of them new—to make this real.
Sam Bowman@sleepinyourhat
2025-04-23More in my new essay on the Anthropic alignment blog here:
https://alignment.anthropic.com/2025/bumpers/ (Putting up Bumpers)