Curved Inference II: Sleeper Agent Geometry - Extending Interpretability Beyond Probes
Anthropic's research explores new ways to detect deceptive alignment in models, using geometric analysis and naturalistic contexts. It's a must-read for those interested in model interpretability and safety.