Qualifying Exam

Qualifying Exam

Towards Reliable and Interpretable Multimodal Intelligence

 

Download as iCal file

Thursday, May 21, 2026, 02:00pm - 03:00pm

 

Speaker: Liwei Che

Bio

Location : CoRE 305

Committee

Professor Vladimir Pavlovic

Assistant Professor Ruixiang Tang

Assistant Professor Chengzhi Mao

Professor Zheng Zhang

Event Type: Qualifying Exam

Abstract: Large vision-language models (LVLMs) have demonstrated strong multimodal capabilities, yet the internal mechanisms underlying their visual reasoning remain poorly understood, limiting both their reliability and controllability. Our research investigates LVLM visual reasoning from a mechanistic interpretability perspective, with a focus on two representative phenomena: object hallucination and visual counting. First, we show that object hallucinations are closely related to a small subset of image tokens that receive disproportionately high attention during generation. Based on this observation, we identify Hallucinatory Image Tokens (HITs) and propose EAZY, a training-free framework that detects and mitigates hallucinated objects by zeroing out token-level visual causes. Second, we study counting as a minimal yet revealing probe of visual reasoning, and introduce Visual Activation Patching and HeadLens to trace how visual information is grounded, routed across modalities, and aggregated into numerical predictions. This analysis reveals a structured counting circuit composed of visual grounding, cross-modal routing, counting aggregation, and awareness heads. Building on these findings, we develop lightweight interventions that improve both counting robustness and broader visual reasoning. Across experiments, our results show that LVLM failures and capabilities are both governed by sparse, structured, and interpretable internal mechanisms. These findings provide a unified account of how LVLMs process visual information, explain why they succeed or fail in reasoning-intensive settings, and demonstrate that mechanistic insights can directly enable more reliable and more capable multimodal models.

Organization

Contact  Professor Vladimir Pavlovic

Join Zoom Meeting
https://rutgers.zoom.us/j/99544402131?pwd=f0peVj0froNKABgb7xLRshlqpUiUWm.1

Join by SIP


Meeting ID: 995 4440 2131
Passcode: 203433