Qualifying Exam

Qualifying Exam

Towards Efficient Multi-LLM Collaborative Debate via Reinforcement Learning

 

Download as iCal file

Wednesday, May 13, 2026, 01:00pm - 02:00pm

 

Speaker: Ruosong Ye

Bio

Location : CBIM 22

Committee

Professor Dimitris Metaxas

Assistant Professor Hongyi Wang

Associate Professor Konstantinos Michmizos

Assistant Professor Sharon Levy

Event Type: Qualifying Exam

Abstract: LLM-based Multi-Agent Debate (MAD), a test-time scaling method featuring "cognitive inflection points," is a prominent research area. By leveraging the complementary knowledge and reasoning of diverse LLMs, MAD often outperforms single models in complex tasks. However, it faces three key challenges: (1) the necessity of multi-round debate is controversial, as it frequently fails to surpass communication-free collaboration; (2) error propagation and accumulation can lead to a single agent misleading the collective; and (3) redundant communication flows increase computational overhead. The aforementioned drawbacks of MAD highlight the critical need for pruning communication flows and an effective pruning algorithm should enhance collective reasoning while reducing token overhead. However, existing research suffers from several limitations: (1) a lack of rigorous comparison against strong baselines like communication-free frameworks, often coupled with overoptimistic estimates of token efficiency ratios; (2) either proposes static optimization frameworks or relies merely on simple LLM role-based profiles for dynamic pruning, failing to fully leverage the specific details within multi-round debates; and (3) a lack of RL rewards and algorithms specifically tailored for flow pruning in multi-round scenarios. Consequently, current frameworks fail to demonstrate the inherent superiority of debate-based collaboration, and suboptimal RL pipelines leave the potential of MAD largely untapped. To address the above challenges, we propose a lightweight pruning model, trained via an innovative multi-round debate reward. Experimental results show that our framework achieves SOTA performance on reasoning benchmarks such as MATH-500, while significantly reducing token costs under fair and rigorous evaluation.

Organization

Contact  Professor Dimitris Metaxas

Zoom Link: https://rutgers.zoom.us/my/ry233?pwd=SU1MRVV3ZVh2K29aWkVEY2F1WXVZQT09