模型

Baseten 与 GoodfireAI 联合推动开源模型安全与可解释性研究

Happy to start collaborating with @baselabs from @baseten and @GoodfireAI to push safety and interpr...

精选理由

Baseten 和 GoodfireAI 联手,把安全研究直接加到 Baseten 的推理里,让客户用起来更安全,值得看看。

Baseten 与 GoodfireAI 联合宣布将推动开源模型的安全与可解释性研究。他们计划将安全研究直接集成到 Baseten 的推理基础设施中,包括训练模型遵循显式政策、实时检测失败并连接到干预控制。此举旨在为所有客户建立安全与监控标准。

原文 · Thomas Wolf

Happy to start collaborating with @baselabs from @baseten and @GoodfireAI to push safety and interpr...

Happy to start collaborating with @baselabs from @baseten and @GoodfireAI to push safety and interpretability for open models. As open-source models catch up in performance to frontier models, the community at large have a great opportunity to establish a common practice for effective safety, security and interpretability research. Some past discussions confused “open” with “unmonitored”. On the contrary, an open model provides much more tooling and visibility for safety research from a broader audience, which has led to a lot of the safety and security techniques we use today. Linux is a great example of this: open-source and secure deployment are not only compatible but heavily intertwined, as a properly secure system needs an extensive feedback loop of finding and fixing vulnerabilities. We're excited to share more soon. Charlie O'Neill @oneill_c Safety is not just for closed models. The closed frontier labs are a canary in the coal mine for what is coming at scale. They give us a glimpse into the future and a window to harden our systems and prepare for abundant intelligence, with all the risks that come along with it. The OpenAI agent swarm attack on Hugging Face is the kind of failure we need to prepare for as open-source models catch up. The providers serving those models (such as Baseten) have a big role in establishing what safety and monitoring standards look like. We’re proud to be taking the lead on this at @baselabs with our collaborators @huggingface and @GoodfireAI . We’re developing safety research in the open and building it directly into Baseten’s inference infrastructure, with the aim of making it available to all our customers. This includes training models to follow explicit policies, detecting failures at runtime, and connecting those signals to controls that can intervene. We invite others in the open-source ecosystem to join us in building the tools and standards we’ll all need. 🔗 View Quoted Tweet 💬 2 🔄 2 ❤️ 31 👀 3117 📊 5 ⚡