AI模型精选

Gemini 3.8 Flash在文档解析基准上表现不佳

Gemini 3.8 Flash might be maxxed on other benchmarks, but not quite on document parsing. We benchma...

精选理由

Gemini 3.8 Flash在文档解析上反而比3.7版本退步,专家称其过度拟合公共基准。

AI 摘要

Gemini 3.8 Flash在ParseBench基准测试中表现与3.7和3.6 Flash版本相似。它在表格和内容忠实度上略优于3.7 Flash,但在图表和语义格式化上表现较差。该模型在隐藏问题和数据分析方面出现倒退,似乎过度拟合公共基准。

原文 · Jerry Liu

Gemini 3.8 Flash might be maxxed on other benchmarks, but not quite on document parsing. We benchma...

Gemini 3.8 Flash might be maxxed on other benchmarks, but not quite on document parsing. We benchmarked it on ParseBench. It does slightly better on tables + content faithfulness compared to Gemini 3.7 Flash, but a little worse on charts and semantic formatting. The overall scores are similar to those of 3.7 and 3.6 Flash. We definitely need to evolve the benchmark towards higher difficulty documents. In the meantime though, parsing real-world docs remains a real challenge for frontier models! Bindu Reddy @bindureddy 🚨 Gemini 3.8 Flash Is Bench Maxxed And Is A Regression on 3.7 Flash 🫢 - 3.8 is worse on our benchmark than 3.7 - This version appears to be overfit to public benchmarks It does worse on our hidden questions and regresses in data analysis. Edit - That said, Gemini Flash 3.7 is an AWESOME model. Google should stop with more Flash releases and drop Gemini 4.0 🔗 View Quoted Tweet 💬 9 🔄 0 ❤️ 11 👀 3083 📊 9 ⚡