🤖 AI 资讯

· ·
← 返回列表

ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains

arXiv cs.CL2026-09-23 04:00:00AI应用,多模态,Agent智能体,推理思考,微调蒸馏,模型评测,提示工程,招聘HR,论文原文 ↗

arXiv:2609.25055v1 Announce Type: new

Abstract: In this report we present results of the ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains. This competition aimed to advance research in document understanding through the task of Visual Question Answering (VQA). Building upon previous DocVQA benchmarks, this competition introduces challenging reasoning questions over a diverse collection of documents spanning eight domains, including business reports, scientific papers, slides, posters, maps, comics, infographics, and engineering drawings. The competition concluded with 20 valid submissions from 8 teams spanning zero-shot VLMs, OCR and parser-augmented pipelines, agentic retrieval systems, multi-agent ensembles, and fine-tuned multimodal models. The results show that the strongest systems move beyond single-pass prompting and instead rely on structured evidence extraction, retrieval, verification, and orchestration across multiple components.