产品多源确认精选73°

RSIGym提升研究代理性能

精选理由

Evolvent AI把RSI变成系统工程问题,Opus 5在RSIGym上性能翻倍,还发布了完整代码和实验数据。

Evolvent AI发布RSIGym平台,为研究代理提供训练、推理、评估和沙盒服务。使用Opus 5作为研究代理,系统在SWE-bench Verified基准上从17.67%提升至50.33%。平台还引入RSI-Index指标,衡量前沿代理与目标模型的协同改进效果。

原文 · elvis

Recommended read. And I agree that the RSI is also a systems engineering problem. Self-improving agents need better research environments. RSIGym gives a research agent training, inference, evals, and sandboxes as services it can call. The agent spends its budget on experiments instead of rebuilding infra. With Opus 5 as the researcher, the improved system went from 17.67% to 50.33% on SWE-bench Verified. Also cool to see a way to measure the quality of co-evolution between harnesses and models, which is how full-stack AI companies stay on the frontier. Fanqing @FanqingMengAI 💡Our view: RSI is a systems engineering problem, not just a model problem. Progress depends on the environment a research agent works in: what resources it can call, what it can change, and how it runs experiments. That environment should reflect real production workflows and be easy for agents to use. At @Evolvent_AI , we built RSIGym around Everything as a Service. Agents can call established research services and iterate on data, training settings, and harness code. We also introduce RSI-Index to measure how well frontier agents jointly improve a target model’s weights and harness. Across 6 research agents and 5 benchmarks, Opus 5 leads at 0.4809. 🧵 We’re releasing the code, experiment configurations, research trajectories, and evaluation logs. Try your own improvement methods with RSIGym. 🌐 Websit rsi-index.ai UN5 💻 GitH github.com/evolvent-ai/RS… y9OH 📄 ar arxiv.org/abs/2610.10310 y6fIV Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 3 🔄 2 ❤️ 9 👀 1948 📊 4 ⚡