ARC Prize 与 Snorkel 在 SF Tech Week 举办首届基准测试生态晚宴
ARC Prize 拉上 Snorkel 办晚宴,40 个搞基准测试的人聚一起,连 METR、Epoch、Artificial Analysis 都来了,聊怎么把模型评测做标准化。
ARC Prize 联合 Snorkel AI 在 SF Tech Week 举办首届 Frontier AI Benchmarking Ecosystem Dinner。活动邀请约 40 位来自基准测试组织、产业界、学术界和政府的专家,讨论 AI 基准测试的设计框架。METR、Epoch、Vals、Artificial Analysis 以及 Harvard、Stanford 的相关人员将出席。背景是 Google DeepMind、OpenAI、Anthropic 的负责人和政策制定者都在呼吁更严格、标准化的模型评估。
We're hosting our first Frontier AI Benchmarking Ecosystem Dinner with @SnorkelAI during SF Tech Week
The conversation around independent model testing has reached a tipping point, with policymakers and leaders from Google DeepMind, OpenAI, and Anthropic calling for more rigorous and standardized evaluations.
We’re bringing together 40 experts from nonprofit benchmarking organizations, industry, academia, and government to discuss frameworks for AI benchmark design and how we can work together to build a more scalable and transparent ecosystem
Looking forward to seeing friends from METR, Epoch, Vals, Artificial Analysis, Harvard, Stanford and more to discuss the future of this important domain