Resources
Reading list
Papers, courses, and tools that have shaped how I think about AI safety. I curate and annotate the list myself, and I try to be honest about what actually moved my understanding. I update it as I work through the research agenda.
New to AI safety? Start with Concrete Problems (2016) for grounding, then Risks from Learned Optimization (2019) for the conceptual frame. BlueDot's AI Safety Fundamentals course is the best structured on-ramp I've found. Full disclosure: I facilitate for BlueDot and my transition year is funded by their grant, so weigh the recommendation accordingly.
Foundational
- 2016
- 2019
- 2017
- 2022
Open-weight & fine-tuning safety
My primary research niche is safety properties that must survive fine-tuning, quantization, and weight release. These papers are the empirical bedrock.
- 2023
- 2023
- 2023
- 2024
Interpretability
- 2022
- 2023
- tool
Multi-agent & compositional alignment
- 2026
- 2023
Evaluation & tools
- 2024
- tool
Courses
- course
- course
Where the discourse lives
- Alignment Forum is the primary venue for technical alignment research and discussion. It's the place to read new work before it becomes a paper.
- LessWrong — broader rationalist / AGI discourse. Uneven, and the best posts are worth the dig.
- Anthropic Research — interpretability, evals, and alignment, published in unusual technical detail.
- Redwood Research — adversarial robustness and scalable oversight from a small, focused team.