Cultural Heritage and Inclusive Societies

Area/s: Cultural Heritage and Inclusive Societies

Organization: Computer Vision Center

Research theme code : RLA-CVC-06

Brief Theme Description:

The field of Document Understanding encompasses the techniques and methods that allow machines to analyse and interpret documents through vision. It stands at the frontier between computer vision and natural language processing, bridging the gap between how machines perceive visual content and comprehend language.
Recent advancements in Multimodal Large Language Models (LLMs) have given a significant boost to this field. These developments have opened up new possibilities, enabling the creation of foundational models for document understanding that are more powerful and versatile than ever before.
Our research seeks to build on this foundation to advance holistic Document Understanding (DU) and address areas that remain underexplored by current multimodal LLMs. Specifically, we aim to tackle these key challenges: (i) enabling models to handle large-scale inputs, such as multi-page, multi-document, and high-resolution large-scale documents; (ii) provide models with the capacity for advanced reasoning on heterogeneous content types and multiple sources of evidence, integrating specialized tools in an agentic framework (iii) developing interpretable and explainable DU models; and (iii) ensuring that models provide privacy guarantees regarding sensitive information used in their training data.

Available Infrastructures

GPU clusters (A40, L40).

Possible Secondments

Yooz (France), Letxbe (France), Rossum (Czech Republic), Snowflake (Poland), CISPA (Germany), Univ. of Helsinki (Finland), IIIT Hyderabad (India).

Keywords

Multimodal Large Language Models; Document Visual Question Answering; Trustworthy AI; Privacy; Robustness; Explainability.

Principal Investigators

Dimosthenis Karatzas and Ernest Valveny lead research on multimodal document collection and trustworthy AI.
Email: dimos@cvc.uab.es ; ernest@cvc.uab.es
Web: http://vlr.ai/