Resources

Reading list

Papers, courses, and tools that have shaped how I think about AI safety. I curate and annotate the list myself, and I try to be honest about what actually moved my understanding. I update it as I work through the research agenda.

New to AI safety? Start with Concrete Problems (2016) for grounding, then Risks from Learned Optimization (2019) for the conceptual frame. BlueDot's AI Safety Fundamentals course is the best structured on-ramp I've found. Full disclosure: I facilitate for BlueDot and my transition year is funded by their grant, so weigh the recommendation accordingly.

Foundational

Open-weight & fine-tuning safety

My primary research niche is safety properties that must survive fine-tuning, quantization, and weight release. These papers are the empirical bedrock.

Interpretability

Multi-agent & compositional alignment

Evaluation & tools

Courses

Where the discourse lives