Projects

Projects

This is where the research, the teaching, and the building actually happen.

Inoculation against model poisoning — SAIGE

A research fellowship with Safe AI Germany, where I test data-level antidote datasets against emergent misalignment on small open-weight models.

Inverse scaling & incoherence

A mechanistic study of why more test-time reasoning sometimes makes models worse. I'm doing intra-trace causal analysis with steering-vector interventions, and writing up what I find as I go.

Multi-agent alignment (HCII 2026)

A peer-reviewed study showing that single- and multi-agent LLM systems behave measurably differently on alignment. I read it as early evidence for compositional misalignment.

Paper knowledge base

A large, interlinked set of paper notes and research logs across AI safety, interpretability, and ML. It's the reading and thinking space behind the writing here.

BlueDot Impact — facilitation

Facilitating technical AI safety cohorts, helping mid-career engineers cross from AI literacy into real contribution. It's the pipeline gap I went through myself.

An autonomous agent system

I build and run a self-improving personal AI agent with persistent memory and standing skills. It's hands-on practice with the exact kind of agentic system whose safety I study.