Brief Theme Description:
The field of Document Understanding encompasses the techniques and methods that allow machines to analyse and interpret documents through vision. It stands at the frontier between computer vision and natural language processing, bridging the gap between how machines perceive visual content and comprehend language.
Recent advancements in Multimodal Large Language Models (LLMs) have given a significant boost to this field. These developments have opened up new possibilities, enabling the creation of foundational models for document understanding that are more powerful and versatile than ever before.
Our research seeks to build on this foundation to advance holistic Document Understanding (DU) and address areas that remain underexplored by current multimodal LLMs. Specifically, we aim to tackle these key challenges: (i) enabling models to handle large-scale inputs, such as multi-page, multi-document, and high-resolution large-scale documents; (ii) provide models with the capacity for advanced reasoning on heterogeneous content types and multiple sources of evidence, integrating specialized tools in an agentic framework (iii) developing interpretable and explainable DU models; and (iii) ensuring that models provide privacy guarantees regarding sensitive information used in their training data.
Available Infrastructures
GPU clusters (A40, L40).
Possible Secondments
Yooz (France), Letxbe (France), Rossum (Czech Republic), Snowflake (Poland), CISPA (Germany), Univ. of Helsinki (Finland), IIIT Hyderabad (India).
Keywords
Multimodal Large Language Models; Document Visual Question Answering; Trustworthy AI; Privacy; Robustness; Explainability.
Principal Investigators
Dimosthenis Karatzas and Ernest Valveny lead research on multimodal document collection and trustworthy AI.
Email: dimos@cvc.uab.es ; ernest@cvc.uab.es
Web: http://vlr.ai/

