Home
About
Blog
Blog
Sep 2026
Character training can mitigate reward hacking, but can also make it harder to detect
May 2026
Bodhisattva as an Alignment Target
(slides)
Apr 2026
Buddhist Wisdom and the Challenge of AI Emotions
(slides)
Apr 2026
Buddhist Wisdom and the Challenge of AI Emotions
(video)
Dec 2025
Automated Evaluation Builder
Dec 2025
My Experience with Automating AI Evaluation Production
Old blog
Jan 2025
Automating AI Control Evaluations - Exploration Summary
May 2024
Ideal Responsible Scaling Policies
May 2024
Policy for Mitigating Catastrophic Risks from AI
Mar 2024
My Math PhD - Summary
Mar 2024
Explaining the AI Alignment Problem to Tibetan Buddhist Monks
Mar 2024
Anomalous Concept Detection for Detecting Hidden Cognition
Feb 2024
Hidden Cognition Detection Methods and Benchmarks
Feb 2024
Notes on Internal Objectives in Toy Models of Agents
Oct 2023
Internal Target Information for AI Oversight
Sep 2023
High-level interpretability: detecting an AI's objectives
May 2023
Aligned AI via monitoring objectives in AutoGPT-like systems
Nov 2022
Auditing games for high-level interpretability
Jul 2022
Deception?! I ain’t got time for that!
???
Kaczynski’s Self-Propagating Systems Theory and the Future of Humanity