HyperBrowseComp 基准发布:13 种语言考验浏览智能体
HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents
一个专门折磨浏览智能体的新基准:423 道题、13 种语言,还要看视频和扫描件找答案,各家模型都被拉出来测了。
HyperBrowseComp 是一个多语言、多模态的网页浏览基准,包含 423 道人工编写并经人工校验的题目,覆盖 13 种语言,均由母语者撰写。题目要求模型定位隐晦证据、串联多步线索,或检查视频、扫描文档、图片、地图等异构来源。出题时先用无联网模型过滤掉可凭参数知识直接回答的简单题。评测采用统一的智能体协议,比较了多个模型在原生搜索和共享外部检索工具下的表现,并抽样做了人类对照测试。
HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents
We introduce HyperBrowseComp, a multilingual and multimodal browsing benchmark comprising 423 manually authored and human-validated questions across 13 languages, written by native or highly proficient speakers. Questions are designed to be extremely challenging. Each question targets a concise, publicly verifiable answer whose discovery requires locating obscure evidence, following multi-step clue chains, or inspecting heterogeneous sources such as videos, scanned documents, images, or maps. Easier questions are filtered out by evaluating them with models without internet access to reduce the likelihood that they can be answered with parametric knowledge alone. We evaluate several models using provider-native search and a shared external retrieval harness under a common agent protocol. To contextualize model performance and effort, we also conduct a human evaluation on a sample of the questions. HyperBrowseComp provides a challenging testbed for persistent information seeking across languages and evidence modalities, with difficulty arising from discovering and connecting evidence on the open web.