Inoculation against model poisoning — SAIGE
A research fellowship with Safe AI Germany, where I test data-level antidote datasets against emergent misalignment on small open-weight models.
Projects
This is where the research, the teaching, and the building actually happen.
A research fellowship with Safe AI Germany, where I test data-level antidote datasets against emergent misalignment on small open-weight models.
A mechanistic study of why more test-time reasoning sometimes makes models worse. I'm doing intra-trace causal analysis with steering-vector interventions, and writing up what I find as I go.
A peer-reviewed study showing that single- and multi-agent LLM systems behave measurably differently on alignment. I read it as early evidence for compositional misalignment.
A large, interlinked set of paper notes and research logs across AI safety, interpretability, and ML. It's the reading and thinking space behind the writing here.
Facilitating technical AI safety cohorts, helping mid-career engineers cross from AI literacy into real contribution. It's the pipeline gap I went through myself.
I build and run a self-improving personal AI agent with persistent memory and standing skills. It's hands-on practice with the exact kind of agentic system whose safety I study.