产品精选73°

LlamaIndex发布专用表单解析模型

精选理由

LlamaIndex发布了专门解析复杂表单的模型,能处理手写和100+字段,比通用VLM更精准便宜。

LlamaIndex构建了最先进的表单阅读模型,专为处理复杂表单文档设计。该模型能准确识别100+字段的表单,包括手写和复选框,并精确定位每个值来源。LlamaParse提供低成本解决方案,处理W-2、1040等税务表单准确率接近100%。

图片来源 · Jerry Liu
原文 · Jerry Liu

We built state-of-the-art models for reading forms 📋 Form documents have the following properties that trip up VLMs: ✅ They carry much more structure than can be represented in standard markdown. You need consistent types for checkboxes, textboxes, labels, signature fields. ✅ They can be extremely complicated (forms can be scanned, there can handwriting scribbles, some forms are dense with ~100+ fields) but accuracy requirements need to be close to 100% ✅ Any form parser requires accurate grounding and attribution. Not only should you extract the values, but you should also be able to precisely locate where each value came from in the source doc ✅ Any form needs to be not only accurate, but cheap/fast We've done a deep-dive into what it takes to build a form parser in this blog post: llamaindex.ai/blog/why-vlms-… d If you want to try out our form models, check out LlamaParse: cloud.llamaindex.ai 8 Your browser does not support the video tag. 🔗 View on Twitter LlamaIndex 🦙 @llama_index Parsing forms still trips up the latest frontier VLMs. A form isn't text on a page. It's a set of fields, grouped into sections, each tied to a specific box. That's why forms need purpose-built parsing, not a bigger general model: ✅️ Detect every field, not just the obvious ones ✅️ Keep the hierarchy of sections and fields ✅️ Tie every value to the exact box it came from ✅️ Understand handwriting and checkmarks Our latest blogpost breaks down where VLMs fail on real W-2s, 1040s, W-9s and scanned W-4s, and a custom cookbook for LlamaParse to handle them at a fraction of the cost 👇 llamaindex.ai/blog/why-vlms-… y 🔗 View Quoted Tweet 💬 2 🔄 0 ❤️ 8 👀 1227 📊 4 ⚡