Joseph E. Gonzalez

Machine Learning + Computer Systems · UC Berkeley

Joseph E. Gonzalez

About Joseph E. Gonzalez

I am a professor in EECS at UC Berkeley, working at the intersection of machine learning and computer systems. My research asks how new systems abstractions can enable new forms of machine learning and AI.

Over the past fifteen years, my students and collaborators have built systems spanning large-scale graph learning, distributed machine learning, model serving and tuning, LLM inference and evaluation, and AI agents. This work includes GraphLab, PowerGraph, GraphX, Clipper, Ray Serve, Ray Tune, Alpa, vLLM, SGLang, Chatbot Arena, LLM-as-a-Judge, Gorilla, the Berkeley Function Calling Leaderboard (BFCL), MemGPT, SkyRL and LEANN, among other projects.

A recurring theme is building not only new algorithms but also the systems and research communities that make new ideas practical and widely used. Many of these systems were led by PhD students I advised, who have gone on to academic positions, AI research labs, and companies they founded.

One question has shaped my agenda: If we solve today's hot problems, what new problems will those solutions create? In 2008, when most of the field was developing new models, I bet that scaling learning to much larger datasets and more parallel compute would become the challenge, and began working on ML systems. In 2015, as the field focused on scaling training, I bet that serving large models and adapting them to new contexts would come next, and began working on inference systems.

Today I am betting that many of the hard problems in AI agents will be solved. My group asks how systems must change for thousands of highly capable agents, how people can steer and negotiate through them, and how models can move from learning from data to learning from teachers. We also work on post-training and new neural architectures. More on these bets →

Founding member of the Sky Computing Lab, LMSYS and RISELab; member of BAIR.
Research story · Bio · CV · Google Scholar · Students' podcast on PhD life

Office: 4165 Gateway.
Office hours: Mondays, 3:00–4:45 PM PST, in Gateway B1013.

Research Through the Years

A single line of inquiry: each era asks what systems a new kind of machine learning needs, and each builds on the ideas and people of the one before.

  1. 2008–2014

    Learning at Scale

    Machine-learning algorithms have structure that systems can exploit.

    GraphLab and PowerGraph, from doctoral work at CMU · GraphX, now part of Apache Spark

  2. 2015–2021

    ML Infrastructure

    Training a model is only the beginning; the rest of the ML lifecycle is a systems problem.

    Clipper (Daniel Crankshaw) · Ray Serve (Simon Mo) · Ray Tune (Richard Liaw)

  3. 2021–2024

    Foundation Models

    Large models made inference efficiency and evaluation central research problems.

    Alpa (Lianmin Zheng) · vLLM · SGLang (Lianmin Zheng, Ying Sheng) · Chatbot Arena and LLM-as-a-Judge · Gorilla (Shishir Patil, Tianjun Zhang)

  4. 2023–present

    Agents, Post-Training and New Architectures

    If many hard problems in agents are solved, systems, human oversight and learning itself will have to change.

    Agent-first systems · humans steering agent teams · learning from teachers · MemGPT (Charles Packer, Sarah Wooders, Kevin Lin) · SkyRL · LEANN · new architectures (MRNN, sparse attention)

Read the full research story, including the people behind each project →

Alumni

Former students and postdocs, with the thesis and projects from their time in the group and selected subsequent roles, as listed in my CV. Dates indicate advising periods, not necessarily full degree enrollment or completion.

Recent papers from the group. Foundational papers for each project are linked from the research story, and the full list is on Google Scholar and in the CV.

Autellix: An Efficient Serving Engine for LLM Agents as General Programs

AuthorsMichael Luo, Xiaoxiang Shi, Colin Cai, Tianjun Zhang, Justin Wong, Yichuan Wang, Chi Wang, Yanping Huang, Zhifeng Chen, Joseph E. Gonzalez, Ion Stoica
NSDI 2026

vCache: Verified Semantic Prompt Caching

AuthorsLuis Gaspar Schroeder, Aditya Desai, Alejandro Cuadron, Kyle Chu, Shu Liu, Mark Zhao, Stephan Krusche, Alfons Kemper, Matei Zaharia, Joseph E. Gonzalez
ICLR 2026

Search Arena: Analyzing Search-Augmented LLMs

AuthorsMihran Miroyan, Tsung-Han Wu, Logan King, Tianle Li, Jiayi Pan, Xinyan Hu, Wei-Lin Chiang, Anastasios N. Angelopoulos, Trevor Darrell, Narges Norouzi, Joseph E. Gonzalez
ICLR 2026

Characterizing Agents in Production

AuthorsMelissa Z. Pan, Negar Arabzadeh, Riccardo Cogo, Yuxuan Zhu, Alexander Xiong, Lakshya A. Agrawal, Huanzhi Mao, Emma Shen, Sid Pallerla, Liana Patel, Shu Liu, Tianneng Shi, Xiaoyuan Liu, Jared Quincy Davis, Emmanuele Lacavalla, Alessandro Basile, Shuyi Yang, Paul Castro, Daniel Kang, Koushik Sen, Dawn Song, Joseph E. Gonzalez, Ion Stoica, Matei Zaharia, Marquita Ellis
ICML 2026

Supporting Our AI Overlords: Redesigning Data Systems to Be Agent-First

AuthorsShu Liu, Soujanya Ponnapalli, Shreya Shankar, Sepanta Zeighami, Alan Zhu, Shubham Agarwal, Ruiqi Chen, Samion Suwito, Shuo Yuan, Ion Stoica, Matei Zaharia, Alvin Cheung, Natacha Crooks, Joseph E. Gonzalez, Aditya G. Parameswaran
CIDR 2026

vAttention: Verified Sparse Attention via Sampling

AuthorsAditya Desai, Kunal Kumar Agrawal, Sehoon Yang, Alejandro Cuadron, Luis Gaspar Schroeder, Matei Zaharia, Joseph E. Gonzalez, Ion Stoica
ICLR 2026

FrontierCS: Evolving Challenges for Evolving Intelligence

AuthorsQiuyang Mang, Wenhao Chai, Zhifei Li, Huanzhi Mao, Shang Zhou, Alexander Du, Hanchen Li, Shu Liu, Edwin Chen, Yichuan Wang, Xieting Chu, Zerui Cheng, Yuan Xu, Tian Xia, Zirui Wang, Tianneng Shi, Jianzhu Yao, Yilong Zhao, Qizheng Zhang, Charlie Ruan, Zeyu Shen, Kaiyuan Liu, Zhaoyang Hong, Alex Gu, Ziyi Zhang, Runyuan He, Dong Xing, Zerui Li, Zirong Zeng, Yige Jiang, Lufeng Cheng, Ziyi Zhao, Youran Sun, Suyang Zhong, Junpeng Wang, Donglin Li, Wenyuan Huang, Jialiang Gu, Wesley Zheng, Wangmeiyu Zhang, Ruyi Ji, Xuechang Tu, Zihan Zheng, Zhaozi Wang, Zexing Chen, Jingbang Chen, Jialu Zhang, Aleksandra Korolova, Peter Henderson, Pramod Viswanath, Vijay Ganesh, Saining Xie, Zhuang Liu, Dawn Song, Sewon Min, Ion Stoica, Joseph E. Gonzalez, Jingbo Shang, Alvin Cheung
ICML 2026

How to Train Your Advisor: Steering Black-Box LLMs with Advisor Models

AuthorsParth Asawa, Alan Zhu, Abigail O'Neill, Matei Zaharia, Alexandros G. Dimakis, Joseph E. Gonzalez
ICML 2026

Are Large Reasoning Models Interruptible?

AuthorsTsung-Han Wu, Mihran Miroyan, David M. Chan, Trevor Darrell, Narges Norouzi, Joseph E. Gonzalez
ICML 2026

ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving

AuthorsYifan Qiao, Shan Yu, Shu Anzai, Haoran Ma, Shuo Yang, Yang Wang, Miryung Kim, Yongji Wu, Yang Zhou, Jiarong Xing, Joseph E. Gonzalez, Ion Stoica, Harry Xu
ICML 2026

MFCL Audio: An Audio Function Calling Evaluation for Large Language Models

AuthorsHuanzhi Mao, Aditya Ghai, Imra Dawoodani, Tony Ginart, Shishir G. Patil, John Emmons, Joseph E. Gonzalez
ICML 2026

SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse-Linear Attention

AuthorsJintao Zhang, Haoxu Wang, Kai Jiang, Shuo Yang, Kaiwen Zheng, Haocheng Xi, Ziteng Wang, Hongzhou Zhu, Min Zhao, Ion Stoica, Joseph E. Gonzalez, Jianfei Chen, Jun Zhu
ICLR 2026

Recent Preprints · 2026

  • 2027 · Upcoming: General Chair, Conference on Machine Learning and Systems (MLSys).
  • May 2026: Invited talk at the Samsung Research America AI Summit, “Finding a New North Star for AI Research in the Post AGI Era.”
  • April 2026: Invited faculty talk at the Amazon AI PhD Fellowship Summit, “An Algorithm for Innovation.”
  • 2025–2026: CS Division Vice Chair for Graduate Matters, co-chair of the EECS Graduate Matters Committee, and elected member of Berkeley's Academic Senate Divisional Council.
  • Fall 2025: Co-taught CS 189/289A: Introduction to Machine Learning with Narges Norouzi. Taught Data 100: Principles and Techniques of Data Science in Spring and Fall 2025.
  • September 2025: Invited keynote at the Samsung AI Forum.