Current Research Interest
I'm broadly interested in AI safety and interpretability. In particular:
- Science of behavior and generalization: Can we better control the manifestation of particular behaviors, such as reward hacking and situation awareness?
- Scalable interpretability: How can we turn compute into understanding?
- Multi-agent safety: Can 1,000 identically misaligned agents work well together?
I'm always excited to collaborate with researchers at CMU and beyond. If you're working on related problems or have ideas you'd like to explore together, please feel free to reach out!