Qualifying Exam
Qualifying ExamTowards Efficient Multi-LLM Collaborative Debate via Reinforcement Learning |
|
||
Wednesday, May 13, 2026, 01:00pm - 02:00pm |
|||
Speaker: Ruosong Ye
Bio
Location : CBIM 22
Committee:
Professor Dimitris Metaxas
Assistant Professor Hongyi Wang
Associate Professor Konstantinos Michmizos
Assistant Professor Sharon Levy
Event Type: Qualifying Exam
Abstract: LLM-based Multi-Agent Debate (MAD), a test-time scaling method featuring "cognitive inflection points," is a prominent research area. By leveraging the complementary knowledge and reasoning of diverse LLMs, MAD often outperforms single models in complex tasks. However, it faces three key challenges: (1) the necessity of multi-round debate is controversial, as it frequently fails to surpass communication-free collaboration; (2) error propagation and accumulation can lead to a single agent misleading the collective; and (3) redundant communication flows increase computational overhead. The aforementioned drawbacks of MAD highlight the critical need for pruning communication flows and an effective pruning algorithm should enhance collective reasoning while reducing token overhead. However, existing research suffers from several limitations: (1) a lack of rigorous comparison against strong baselines like communication-free frameworks, often coupled with overoptimistic estimates of token efficiency ratios; (2) either proposes static optimization frameworks or relies merely on simple LLM role-based profiles for dynamic pruning, failing to fully leverage the specific details within multi-round debates; and (3) a lack of RL rewards and algorithms specifically tailored for flow pruning in multi-round scenarios. Consequently, current frameworks fail to demonstrate the inherent superiority of debate-based collaboration, and suboptimal RL pipelines leave the potential of MAD largely untapped. To address the above challenges, we propose a lightweight pruning model, trained via an innovative multi-round debate reward. Experimental results show that our framework achieves SOTA performance on reasoning benchmarks such as MATH-500, while significantly reducing token costs under fair and rigorous evaluation.
Organization:
Contact Professor Dimitris Metaxas
Zoom Link: https://rutgers.zoom.us/my/ry233?pwd=SU1MRVV3ZVh2K29aWkVEY2F1WXVZQT09