Preparing for Automated Alignment
Discussing a paper I wrote, on the risky bet of automated alignment
So, a while ago I became curious about Scalable Oversight and AI Control. Both because they’re cool to work on, and because they seem quite important — we’re racing toward ASI, and in the absence of great theoretical breakthroughs they might be the best techniques we’ll have.1
With Jan Leike’s recent optimism, and Opus 4.6 being quite good, it finally feels like the right time to post about a paper I wrote on this.
The question I was asking: how can we tell if automated alignment will fail?
Main problem framing
We will try to automate AI safety work — have AIs solve the “Superalignment problem” — which seems good. But early successes might not leave much margin for error if things suddenly fail.
“Human-level” automated AI researchers aren’t immune to failure. Even absent overt scheming, they may not help us sufficiently in figuring out ASI alignment.
The handover of research from humans to near-human-level AI systems might seem fine at first — but it would be easy to get complacent.
Already, we’re not very good at catching errors in model outputs. These errors will have higher stakes soon. Can we trust models to know how to safely align ASI?
I wrote a paper trying to formalize this argument and suggest some mitigations:
Testing it: an empirical research agenda to get early hints at the limitations of automated alignment
Governance: we need to know when to draw the line. Near-handover might be that line.

Empirical agenda
In the paper (Appendix A), I proposed blindly testing a group of AI safety researchers given faulty outputs — or a faulty model — to see if they’d catch the flaws.
The idea: if researchers don’t find fairly simple flaws, we shouldn’t expect them to find more complex ones easily — even without scheming, even without asking “will this technique generalize to ASI?”
I think the research agenda mostly doesn’t work in its current form, for various reasons. But a similar approach might, and might be worth testing. The most similar attempt was Marks et al. — a really cool paper pitting several interpretability teams against a tampered model.
What’s next?
At this point I’m not seriously working on this project — but I might be open to collaboration if someone has ideas on how to take it forward.
Alternatively, we might skip the proposed experiment and work on mitigations assuming it was successful — e.g., assuming human researchers are quite bad at detecting model output flaws, what’s next? Testing the generalization of LLMs on AI safety work, or something else.
Acknowledgements
Full list in the paper, but special mention to Dan Valentine, Robert Adranga, Leo Zovic, Sheikh Abdur Raheem Ali, Evgenii Opryshko, Yonatan Cale, Yoav Tzfati, and Anson Ho for discussions and feedback on the research agenda.
The paper, research agenda, and post was not endorsed by and does not necessarily reflect the views of the people listed here.
Of course, interpretability is making good progress. But most people working on it don’t expect it to be the main tool that aligns ASI.

