research
Research in Vision and Language AI, Large Language Models, and Conversational AI.
Research Lines
My research mostly spans five interconnected lines. Large Vision-Language Models — developing and evaluating LVLMs capable of complex video understanding, instruction following, and multi-step reasoning. Large Language Models and Adaptation — efficient fine-tuning and steering of LLMs through parameter-efficient methods, as well as the development of European Portuguese-language models such as GlórIA and AMALIA. Conversational and Agentic AI — task-oriented and open-domain dialogue systems, multimodal conversational agents, and agentic systems that plan and act in tool-augmented environments. Trustworthy AI — safety, robustness, and reliability of AI systems, including red-teaming LLMs and guardrails for agentic systems. Multimodal Understanding — Multimodal representation learning and temporal multimodal embedding spaces that jointly model vision and language over time.
Funded Projects
| Project | Role | Funder | Period |
|---|---|---|---|
| Amazon Nova AI Challenge — Trusted Agents (2nd ed.) | Principal Investigator | Amazon (US) | 2026–present |
| AMALIA — National Portuguese LLM | Co-PI | PRR Programme | 2025–present |
| ViewSport — Visual Intelligence for Sport | Co-PI | PT2030 | 2026–present |
| TACT-NAVY — Trustworthy AI Agent for Navy Tasks | Co-PI | FCT Public Admin. | 2025–2026 |
| Amazon Nova AI Challenge — Trusted AI (1st ed.) | Principal Investigator | Amazon (US) | 2024–2025 |
| Trustworthy LLM for Portuguese and English | Co-PI | Google AI Cloud Grant | 2024–2025 |
| Alexa Prize TaskBot Challenge (TWIZ — 1st place) | Co-PI | Amazon (US) | 2022–2023 |
| Alexa Prize TaskBot Challenge (TWIZ) | Co-PI | Amazon (US) | 2021–2022 |
| iFetch — Multimodal Conversational Agents for Fashion | Senior Researcher | CMU-Portugal | 2020–2023 |
PhD Students
| Student | Period | Topic |
|---|---|---|
| Iago Paulo | 2026– | Learning Reinforced Multimodal-Dialog Strategies for Conversational AI Agents |
| Henrique Paz | 2026– | Learning Reinforced Multimodal-Dialog Strategies for Conversational AI Agents |
| João Pereira | 2023– | Extending Spatio-Temporal Action Detection for Video Surveillance |
| Helin Li | 2023– | ML-assisted Design of Micro-nano Devices (KU Leuven, co-advisor) |
| Diogo Silva | 2022– | Towards Human-like Domain-Aware Multimodal Dialog Agents (FCT NOVA + CMU) |
| Diogo Tavares | 2022– | Tracking Dialog State with Self-Attention for Negotiation Dialogs (FCT NOVA + CMU, co-advisor) |
| Rafael Ferreira | 2021– | Learning Reinforced Multimodal-Dialog Strategies for Conversational AI Agents (co-advisor) |
MSc Students
Ongoing (2025/2026)
| Student | Thesis |
|---|---|
| Carolina Diogo | Learning Subtype-Specific microRNA Regulators for Personalized Breast Cancer Therapy |
| José Romano | Learning to Map microRNA Biomarker Signatures to Breast Cancer Subtypes |
| Sebastião Martins | Eliciting Temporal Knowledge in Large Language Models |
| António Almeida | Spatio-Temporal Reasoning in Large Vision-Language Models |
| Máximo Volynets | Cultural Cognition in LLMs — Bridging Language and Cultural Knowledge |
| Vicente Felino | Intelligent and Contextual Overlays for Live Football Broadcasts |
| Mariana Aguiar | Safety of Code-focused Large Language Models Agents |
| James Hertz | Towards Semantically-Bounded Trustworthy Large Language Models |
| Manuel Letras (co-advisor) | Dynamic Visual Token Representations for Efficient Multimodal Reasoning |
Completed (2022–2025)
| Student | Year | Thesis |
|---|---|---|
| Henrique Paz | 2025 | Exploiting Time in Large Vision and Language Models |
| Pedro Domingos | 2025 | Video QA for Surveillance with Large Vision and Language Models |
| Iago Rodrigues | 2025 | A Trustworthy PT-PT Document-grounded Dialog Large Language Model |
| Artur Horal (co-advisor) | 2025 | Adaptive Attack Planning for Red-Teaming Large Language Models |
| Daniel Pina (co-advisor) | 2025 | Structure-aware PT-PT LLM for Retrieval Augmented Generation |
| Miguel Fortuna (co-advisor) | 2025 | Automatic Identification of Fields in Arbitrary Documents Using LLMs |
| Bernardo Calvo | 2024 | Where and When? Large Vision and Language Models for Multimedia Event Extraction |
| Catarina Bento | 2024 | Conditional Deep Generative Models for MEMS Devices Design |
| Luís Tripa | 2024 | Combining Deep Learning and Optimization for MEMS Devices Design |
| Ricardo Barqueira | 2023 | Multimodal On-the-fly News Media Exploration |
| João Pereira | 2023 | Real-time Human Action Localization in the Wild |
| João Arvana | 2023 | Time-aware Question-Answering for the Portuguese Web Archive |
| Daniel Castanho | 2023 | Structuring and Organizing Large Scale Graph Temporal Information |
| Ricardo Valverde | 2023 | Rich Large-Scale Portuguese Language Models from Large Portuguese Corpora |
| Cláudio Bartolomeu | 2022 | Context-Aware Multimodal Embeddings for Interactive News Images Search |
| Pedro Almeida (co-advisor) | 2022 | Turn-Based Temporal Media Web Visualization and Querying for News Images |
| Jonas Rodrigues (co-advisor) | 2022 | Temporally Smoothed Joint Item-Session Space for Session-based Recommendation |
| Frederico Vicente (co-advisor) | 2022 | Talk Commonsense To Me! Enriching Language Models with Commonsense Knowledge |
| Carolina Lopes | 2022 | Visual Question Answer for News Stories |
| Alexandre Correia (co-advisor) | 2022 | Interactive Fashion Search and Recommendation |
| Diogo Tavares (co-advisor) | 2021 | Joint Dialog State Tracking and Slot Filling with the Transformer |
| Diogo Silva (co-advisor) | 2021 | Dialog Generation for Recommendation Agents |
| Rafael Ferreira (co-advisor) | 2020 | Tracking Context in Conversational Search: From Utterances to Neural Embeddings |
| Mariana Ferreira (co-advisor) | 2020 | Knowledge-Driven Answer Generation for Conversational Search |
| Ruslan Padnevych (co-advisor) | 2020 | SmartyFlow: Robust Facial Biometrics for Virtual Identification |