research

Research in Vision and Language AI, Large Language Models, and Conversational AI.

Research Lines

My research mostly spans five interconnected lines. Large Vision-Language Models — developing and evaluating LVLMs capable of complex video understanding, instruction following, and multi-step reasoning. Large Language Models and Adaptation — efficient fine-tuning and steering of LLMs through parameter-efficient methods, as well as the development of European Portuguese-language models such as GlórIA and AMALIA. Conversational and Agentic AI — task-oriented and open-domain dialogue systems, multimodal conversational agents, and agentic systems that plan and act in tool-augmented environments. Trustworthy AI — safety, robustness, and reliability of AI systems, including red-teaming LLMs and guardrails for agentic systems. Multimodal Understanding — Multimodal representation learning and temporal multimodal embedding spaces that jointly model vision and language over time.


Funded Projects

Project Role Funder Period
Amazon Nova AI Challenge — Trusted Agents (2nd ed.) Principal Investigator Amazon (US) 2026–present
AMALIA — National Portuguese LLM Co-PI PRR Programme 2025–present
ViewSport — Visual Intelligence for Sport Co-PI PT2030 2026–present
TACT-NAVY — Trustworthy AI Agent for Navy Tasks Co-PI FCT Public Admin. 2025–2026
Amazon Nova AI Challenge — Trusted AI (1st ed.) Principal Investigator Amazon (US) 2024–2025
Trustworthy LLM for Portuguese and English Co-PI Google AI Cloud Grant 2024–2025
Alexa Prize TaskBot Challenge (TWIZ — 1st place) Co-PI Amazon (US) 2022–2023
Alexa Prize TaskBot Challenge (TWIZ) Co-PI Amazon (US) 2021–2022
iFetch — Multimodal Conversational Agents for Fashion Senior Researcher CMU-Portugal 2020–2023

PhD Students

Student Period Topic
Iago Paulo 2026– Learning Reinforced Multimodal-Dialog Strategies for Conversational AI Agents
Henrique Paz 2026– Learning Reinforced Multimodal-Dialog Strategies for Conversational AI Agents
João Pereira 2023– Extending Spatio-Temporal Action Detection for Video Surveillance
Helin Li 2023– ML-assisted Design of Micro-nano Devices (KU Leuven, co-advisor)
Diogo Silva 2022– Towards Human-like Domain-Aware Multimodal Dialog Agents (FCT NOVA + CMU)
Diogo Tavares 2022– Tracking Dialog State with Self-Attention for Negotiation Dialogs (FCT NOVA + CMU, co-advisor)
Rafael Ferreira 2021– Learning Reinforced Multimodal-Dialog Strategies for Conversational AI Agents (co-advisor)

MSc Students

Ongoing (2025/2026)

Student Thesis
Carolina Diogo Learning Subtype-Specific microRNA Regulators for Personalized Breast Cancer Therapy
José Romano Learning to Map microRNA Biomarker Signatures to Breast Cancer Subtypes
Sebastião Martins Eliciting Temporal Knowledge in Large Language Models
António Almeida Spatio-Temporal Reasoning in Large Vision-Language Models
Máximo Volynets Cultural Cognition in LLMs — Bridging Language and Cultural Knowledge
Vicente Felino Intelligent and Contextual Overlays for Live Football Broadcasts
Mariana Aguiar Safety of Code-focused Large Language Models Agents
James Hertz Towards Semantically-Bounded Trustworthy Large Language Models
Manuel Letras (co-advisor) Dynamic Visual Token Representations for Efficient Multimodal Reasoning

Completed (2022–2025)

Student Year Thesis
Henrique Paz 2025 Exploiting Time in Large Vision and Language Models
Pedro Domingos 2025 Video QA for Surveillance with Large Vision and Language Models
Iago Rodrigues 2025 A Trustworthy PT-PT Document-grounded Dialog Large Language Model
Artur Horal (co-advisor) 2025 Adaptive Attack Planning for Red-Teaming Large Language Models
Daniel Pina (co-advisor) 2025 Structure-aware PT-PT LLM for Retrieval Augmented Generation
Miguel Fortuna (co-advisor) 2025 Automatic Identification of Fields in Arbitrary Documents Using LLMs
Bernardo Calvo 2024 Where and When? Large Vision and Language Models for Multimedia Event Extraction
Catarina Bento 2024 Conditional Deep Generative Models for MEMS Devices Design
Luís Tripa 2024 Combining Deep Learning and Optimization for MEMS Devices Design
Ricardo Barqueira 2023 Multimodal On-the-fly News Media Exploration
João Pereira 2023 Real-time Human Action Localization in the Wild
João Arvana 2023 Time-aware Question-Answering for the Portuguese Web Archive
Daniel Castanho 2023 Structuring and Organizing Large Scale Graph Temporal Information
Ricardo Valverde 2023 Rich Large-Scale Portuguese Language Models from Large Portuguese Corpora
Cláudio Bartolomeu 2022 Context-Aware Multimodal Embeddings for Interactive News Images Search
Pedro Almeida (co-advisor) 2022 Turn-Based Temporal Media Web Visualization and Querying for News Images
Jonas Rodrigues (co-advisor) 2022 Temporally Smoothed Joint Item-Session Space for Session-based Recommendation
Frederico Vicente (co-advisor) 2022 Talk Commonsense To Me! Enriching Language Models with Commonsense Knowledge
Carolina Lopes 2022 Visual Question Answer for News Stories
Alexandre Correia (co-advisor) 2022 Interactive Fashion Search and Recommendation
Diogo Tavares (co-advisor) 2021 Joint Dialog State Tracking and Slot Filling with the Transformer
Diogo Silva (co-advisor) 2021 Dialog Generation for Recommendation Agents
Rafael Ferreira (co-advisor) 2020 Tracking Context in Conversational Search: From Utterances to Neural Embeddings
Mariana Ferreira (co-advisor) 2020 Knowledge-Driven Answer Generation for Conversational Search
Ruslan Padnevych (co-advisor) 2020 SmartyFlow: Robust Facial Biometrics for Virtual Identification