CS Events
Computer Science Department ColloquiumRethinking AI Systems Through Efficient Model Communication |
|
||
Tuesday, March 03, 2026, 10:30am - 12:00pm |
|||
Speaker: Yuhan Liu
Bio
Yuhan Liu is a fifth-year PhD student at the University of Chicago, co-advised by Junchen Jiang and Shan Lu. Her research interest is in building efficient large-scale system and networking support for ML model inference. Her works appeared in top computer system/networking conferences, such as OSDI, SIGCOMM, NSDI. She received MIT EECS rising star, EuroSys best paper award, and UChicago’s Neubauer PhD fellowship for her research. She also leads two open-source projects that build large-scale KV caching layer for efficient LLM inference, and are used in over 30 companies in production, including Google Cloud, Amazon AWS, NVIDIA, IBM etc.
Location : CoRE 301
Committee:
Event Type: Computer Science Department Colloquium
Abstract: For decades, AI models interacted with humans directly through human-readable inputs and outputs (e.g., texts, images). Today, they are used much more ubiquitously and often interact through complex software systems, interacting with other models or software rather than directly with humans. This paradigm shift raises a natural question: can models interact with other models and software using model-native languages?In this talk, I will present my work on facilitating model-native interactions among models and between models and software. To enable more efficient and practical model interactions using model-native states (i.e., KV cache) in LLM systems, my work CacheGen is the first system to share KV cache across different user queries by compressing it into compact bitstreams, and my work DroidSpeak is the first system to share KV cache across different models. My research made real-world impacts via the open-source project, LMCache, widely used in production by top-tier AI companies. Together, these works make LLM inference 5–10x faster than state-of-the-art inference engines. To enable more accurate model-to-software communication, my work ChameleonAPI encodes software code structure into model-native loss functions, allowing models to be retrained for up to 43% higher application-level accuracy in vision applications.
Organization:
Contact Professor Richard Martin
Join Zoom Meeting
https://rutgers.zoom.us/j/2014444359?pwd=WW9ybFNCNVFrUWlycHowSHdNZjhzUT09
Meeting ID: 201 444 4359
Password: 550978
Subscribe to RSS Feed