Who's in the Loop? Interaction and Control in Multimodal AI

Relatore:  Loris Bazzani - University Verona
  martedì 13 ottobre 2026 alle ore 16.30 Sala Verde (solo presenza)

Abstract:
Today's multimodal models are increasingly capable, yet they remain hard to steer: they are trained to answer in one shot, and when they get it wrong the user has no channel to correct, clarify, or guide them. I argue that interaction and control should be first-class design constraints, shaping both how we train models and how we evaluate them, rather than a thin interface layer added after training. My research agenda is framed around three pillars: controllable multimodal data generation for the long tail and for privacy-restricted domains, lightweight adaptation of models into vertical experts, and human-AI interaction and co-design. This talk discusses the first two only briefly and then goes deep into the third, dissecting two recent works. "Interactive Episodic Memory with User Feedback" [1] addresses episodic memory with natural language queries over long egocentric videos (e.g., "Where did I place the mug?"), where real queries are ambiguous yet current methods answer one-shot. We introduce a new task, in which the user refines the prediction through feedback (e.g., "Before this. I'm looking for the big blue mug not the white one"), together with a lightweight training scheme and a plug-and-play module that retrofits existing models with the ability to use feedback. "Benchmarking Interaction, Beyond Policy: a Reproducible Benchmark for Collaborative Instance Object Navigation" [2] turns to embodied agents that must find a target described in natural language and can ask the human for clarification. We propose the first reproducible benchmark for Collaborative Instance Object Navigation, which disentangles two abilities that are usually conflated, navigating and knowing when and what to ask, and propose a dataset of quality-checked reasoning and question-asking traces. With this data we trained a model that is 3x smaller and 70x faster than existing modular methods that generalizes better to unseen objects and environments.

[1] Interactive Episodic Memory with User Feedback. N. Subedi, L. Bazzani, Z. Al-Halah. CVPR 2026. https://arxiv.org/abs/2604.24893
[2] Benchmarking Interaction, Beyond Policy: a Reproducible Benchmark for Collaborative Instance Object Navigation.E. Zorzi, F. Taioli, Y. Wang, M. Cristani, A. Farinelli, A. Castellini, L. Bazzani. BMVC 2026. https://arxiv.org/abs/2604.00265

Bio:
Loris Bazzani is an AI Research Leader with over 15 years of experience, spanning classical computer vision and machine learning to today’s foundation and multimodal generative models. He is currently an adjunct professor at University at Verona and founding research leader at a stealth AI start-up. In his previous role as Principal Scientist at Amazon, he led core research and product efforts across Prime Video, Alexa, and shopping, co-developing architectures for video understanding, vision-language representation, Large Multimodal Models, and diffusion models. His work powered features such as live sports highlights, virtual try-on, interactive product recommendations, and shopping assistants, reaching millions of users and delivering significant business impact. Loris obtained his Ph.D. in Computer Science from the University of Verona (Italy) in 2012, supervised by Prof. Vittorio Murino and Prof. Marco Cristani. He held postdoctoral positions at Dartmouth College with Prof. Lorenzo Torresani, and at the Italian Institute of Technology with Prof. Vittorio Murino. His research has been published in top-tier venues including CVPR, ICCV, ECCV, and ICML, with 50+ publications and patents: https://baz.github.io/

 

Referente
Alessandro Farinelli

Referente esterno
Data pubblicazione
16 settembre 2026

Offerta formativa

Condividi