Inoculation against model poisoning — SAIGE
Research fellowship with Safe AI Germany. Testing data-level antidote datasets against emergent misalignment on small open-weight models.
Projects
This is where the research, the teaching, and the building actually happen.
Research fellowship with Safe AI Germany. Testing data-level antidote datasets against emergent misalignment on small open-weight models.
A mechanistic study of why more test-time reasoning sometimes makes models worse. I'm doing intra-trace causal analysis with steering-vector interventions, and writing up what I find as I go.
A peer-reviewed study showing that single- and multi-agent LLM systems behave measurably differently on alignment. Early evidence for compositional misalignment.
An interlinked digital garden of paper notes and research logs across AI safety, interpretability, and ML. It's the thinking space that feeds the writing here.
Facilitating technical AI safety cohorts, helping mid-career engineers cross from AI literacy into real contribution. It's the pipeline gap I went through myself.
I build and run a self-improving personal AI agent with persistent memory and standing skills. It's hands-on practice with the exact kind of agentic system whose safety I study.