Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
DocLLM is a layout-aware generative language model for multimodal document understanding that extends traditional large language models by incorporating spatial layout information through bounding box coordinates rather than expensive image encoders. It outperforms state-of-the-art LLMs on 14 out of 16 datasets across document intelligence tasks including visual question answering, key information extraction, and document classification.
Parse Score
Sources
aimultiple.com shapes more of what AI says about DocLLM than any other source, at 50% of its citations.
arxiv.org