#model interpretability

1 条
8月26日
论文官方一手精选10:16
Curved Inference II: Sleeper Agent Geometry - Extending Interpretability Beyond Probes

Anthropic's research explores new ways to detect deceptive alignment in models, using geometric analysis and naturalistic contexts. It's a must-read for those interested in model interpretability and safety.

事件专题
官方一手arXiv: Anthropic@Rob Manson8 个信源在谈原文