ACMCV 2025

The Annual Catalan Meeting on Computer Vision (ACMCV) brings together the Computer Vision community of Catalonia, connecting research, talent generation and industry in one day. This meeting aims to strengthen the links among the Catalan Computer Vision actors, to disseminate within the community the most relevant works that have already been published abroad and to allow students from the Master’s Degree in Computer Vision to meet with members of the Catalan computer vision community and prospective employers.

Dates & Venue

📑 Submission deadline | September 9, 2025
✍ Registration deadline | September 16, 2025
🗓 ACMCV 2025 | September 16, 2025
📍 Computer Vision Center & UAB School of Engineering

Program

StartEndSessionLocationMore information
14:0014:45Track1: Msc DefencesEE Rooms / CVC RoomsMore info
14:4515:30Track2: Msc DefencesEE Rooms / CVC RoomsMore info
15:3016:15Track3: Msc DefencesEE Rooms / CVC RoomsMore info
16:1517:00Track4: Msc DefencesEE Rooms / CVC RoomsMore info
17:0017:45Track5: Msc DefencesEE Rooms / CVC RoomsMore info
17:0017:30Accreditation & poster setupCVC Garden 
17:3018:15Industry pitchCVC GardenMore info
18:1518:45Poster session and networkingCVC GardenMore info
18:4519:30Keynote talk: Dr David VázquezCVC GardenMore info
19:3020:00Awards & ClosingCVC Garden 

Keynote talk

Enterprise Visual Understanding with Vision-Language Models: From Documents to Intelligent Agents

Vision-Language Models (VLMs) have demonstrated remarkable progress in natural image understanding and creative generation, yet their performance often falls short on enterprise-critical tasks such as document analysis, chart reasoning, workflow automation, and user interface navigation. In this talk, will be presented recent advances in adapting multimodal foundation models to enterprise applications, with a focus on text-rich visual understanding, document intelligence, and visual content–to–code generation. Also, will be introduced datasets and benchmarks such as BigDocs, BigCharts, StarFlow, and StarVector, designed to push VLMs toward real-world enterprise use cases. It will also be discussed AlignVLM, a robust architecture that bridges visual and textual representations to achieve competitive results on challenging document benchmarks. Finally, it will be highlighted how these models enable the next generation of AI agents—systems capable of reasoning, planning, and acting—by grounding natural language instructions in complex graphical user interfaces. Together, these directions illustrate a path toward enterprise-ready multimodal AI that is accurate, reliable, and adaptable.

Dr. David Vázquez

Staff Research Scientist at Google DeepMind

Resources

Organizers

Organizing Committee

  • Maria Vanrell, CVC & UAB
  • Josep Lladós, CVC & UAB
  • Núria Martínez, CVC
  • Xavier Galvez, CVC
  • Aurora García, CVC
Poster Session Chair:
  • Danna Xue, CVC