Qualifying Exam
Qualifying ExamTowards Reliable and Interpretable Multimodal Intelligence |
|
||
Thursday, May 21, 2026, 02:00pm - 03:00pm |
|||
Speaker: Liwei Che
Bio
Location : CoRE 305
Committee:
Professor Vladimir Pavlovic
Assistant Professor Ruixiang Tang
Assistant Professor Chengzhi Mao
Professor Zheng Zhang
Event Type: Qualifying Exam
Abstract: Large vision-language models (LVLMs) have demonstrated strong multimodal capabilities, yet the internal mechanisms underlying their visual reasoning remain poorly understood, limiting both their reliability and controllability. Our research investigates LVLM visual reasoning from a mechanistic interpretability perspective, with a focus on two representative phenomena: object hallucination and visual counting. First, we show that object hallucinations are closely related to a small subset of image tokens that receive disproportionately high attention during generation. Based on this observation, we identify Hallucinatory Image Tokens (HITs) and propose EAZY, a training-free framework that detects and mitigates hallucinated objects by zeroing out token-level visual causes. Second, we study counting as a minimal yet revealing probe of visual reasoning, and introduce Visual Activation Patching and HeadLens to trace how visual information is grounded, routed across modalities, and aggregated into numerical predictions. This analysis reveals a structured counting circuit composed of visual grounding, cross-modal routing, counting aggregation, and awareness heads. Building on these findings, we develop lightweight interventions that improve both counting robustness and broader visual reasoning. Across experiments, our results show that LVLM failures and capabilities are both governed by sparse, structured, and interpretable internal mechanisms. These findings provide a unified account of how LVLMs process visual information, explain why they succeed or fail in reasoning-intensive settings, and demonstrate that mechanistic insights can directly enable more reliable and more capable multimodal models.
Organization:
Contact Professor Vladimir Pavlovic
Join Zoom Meeting
https://rutgers.zoom.us/j/99544402131?pwd=f0peVj0froNKABgb7xLRshlqpUiUWm.1
Join by SIP
Meeting ID: 995 4440 2131
Passcode: 203433