Fifteen years of research by Joseph E. Gonzalez and his students at UC Berkeley on the systems that make new forms of machine learning possible.
My research asks how new systems abstractions can enable new forms of machine learning and AI. The questions have changed as machine learning has changed, but each era grew out of the one before it, often through the same people. This page tells that story in order: the idea behind each era, the projects that tested it, and the students and collaborators who built them.
If we solve today's hot problems, what new problems will those solutions create?
I have used this question to choose research directions throughout my career. The bet is never that a hot problem stays unsolved. It is that someone solves it, and the solution creates the next problem.
Machine-learning algorithms have structure that systems can exploit.
Many machine-learning algorithms, from belief propagation to matrix factorization, are made of small computations over sparse dependencies, repeated until they converge. General-purpose data systems of the time did not see that structure, so they either ran these algorithms inefficiently or forced researchers to write custom parallel code. My doctoral work at Carnegie Mellon asked whether the graph could serve as the abstraction between the two: machine-learning researchers describe computation on vertices and edges, and the system handles scheduling, consistency and distribution.
GraphLab, PowerGraph and GraphX are three steps in that one line of inquiry. GraphLab introduced the abstraction. PowerGraph re-factored it so that computation could be split across machines for the power-law graphs found in real data. GraphX, built as a postdoc in Berkeley's AMPLab, showed that graph computation could be expressed inside a general dataflow system, unifying tables and graphs in Apache Spark.
GraphLab
A graph-parallel abstraction for machine learning
GraphLab let researchers express iterative ML algorithms as programs on the vertices of a sparse graph, with the system managing asynchronous scheduling and consistency. Distributed GraphLab (VLDB 2012) extended it to the cloud.
People: My doctoral work at Carnegie Mellon, with Yucheng Low, Aapo Kyrola, Danny Bickson and my PhD advisor Carlos Guestrin. The work led to the company Turi (formerly GraphLab), which Apple acquired in 2016.
VLDB 2023 Test of Time Award (Distributed GraphLab)
Real-world graphs have a few very high-degree vertices that overwhelm vertex-centric systems. PowerGraph decomposed vertex programs into gather, apply and scatter phases and partitioned graphs by cutting vertices instead of edges, spreading high-degree vertices across machines.
People: With Yucheng Low, Haijie Gu, Danny Bickson and Carlos Guestrin.
Graph processing in a distributed dataflow framework
GraphX showed that graph-parallel computation can be expressed with ordinary dataflow operators and optimized like database joins, so graphs and tables can share one system. It became the graph library of Apache Spark.
People: Developed in Berkeley's AMPLab with Reynold Xin, Ankur Dave and Daniel Crankshaw, who became one of my first PhD students.
Training a model is only the beginning; the rest of the ML lifecycle is a systems problem.
After I joined the Berkeley faculty, and as part of the RISELab, the focus shifted from training to everything that happens around it. ML workloads were becoming dynamic and heterogeneous: models had to answer live queries within milliseconds, run as pipelines of several models, and be selected from large searches over architectures and hyperparameters. Our view was that these steps needed their own systems abstractions.
The clearest example is prediction serving, a line of work that runs directly into today's LLM inference systems (see the lineage below). Velox and then Clipper, led by my student Daniel Crankshaw, put a general-purpose serving layer between applications and ML frameworks. InferLine extended it to model pipelines. Ray Serve continued these ideas inside the Ray ecosystem. In parallel, Ray Tune, led by my student Richard Liaw, made hyperparameter search a distributed-systems problem. I co-led work on Ray Serve and Ray Tune with students and collaborators, developing ML abstractions for serving and tuning within the Ray ecosystem.
Students in this period also worked on parallel dataframes (Modin), stateful serverless computing (Cloudburst), dynamic neural networks and learned resource management. Their theses are listed under Alumni.
A lineage of serving systems. A continuing line of research from prediction serving to today's LLM inference engines, connected through ideas, students and collaborators.
Clipper put a serving layer between applications and ML frameworks. Caching, adaptive batching and online model selection made predictions fast and accurate regardless of which framework trained the model. It followed Velox (CIDR 2015), and InferLine (SoCC 2020) extended it to provisioning pipelines of models under latency objectives.
Ray Serve carried the prediction-serving ideas of Clipper into Ray: scalable, framework-agnostic deployment of models and multi-model pipelines, written in Python on a general distributed runtime.
People: Ray Serve began as a rewrite of Clipper by Simon Mo, who continued its development within Ray after Daniel Crankshaw graduated.
Distributed hyperparameter tuning and model selection
Tune gave researchers a scalable interface for running and coordinating many training trials, with pluggable search and early-stopping algorithms. Follow-on work (RubberBand, EuroSys 2021) studied cost-efficient tuning on elastic cloud resources.
People: Led by Richard Liaw, then my student, with Eric Liang, Robert Nishihara and Philipp Moritz. RubberBand was led by Ujval Misra with Richard and Lisa Dunlap.
Large models made inference efficiency and evaluation central research problems.
Foundation models raised the systems questions of the previous era at a much larger scale, and added new ones. Alpa, led by my student Lianmin Zheng, automated the parallelization of models too large for a single accelerator, and AlpaServe applied model parallelism to serving. Once LLMs were widely used, inference and evaluation became the central problems. The people who built one of these systems often went on to build the next.
Two inference engines from this period became major infrastructure for the field. vLLM uses operating-system-style paging to manage GPU memory. SGLang treats LLM applications as programs and co-designs a language with a runtime. Chatbot Arena made evaluation an open, live measurement of human preference. Gorilla asked how models can reliably call the APIs and tools they will need to act in the world, a question that leads directly to agents.
vLLM
Efficient LLM serving with PagedAttention
PagedAttention borrows virtual memory and paging from operating systems to store the attention key-value cache in non-contiguous blocks, nearly eliminating memory fragmentation and allowing far larger batches. vLLM grew into one of the most widely used open-source LLM inference engines.
People: Created by Woosuk Kwon and Zhuohan Li with collaborators in the Sky Computing Lab, including my student Lianmin Zheng. Simon Mo, whom I advised, contributed to the open-source project and later co-founded Inferact, the company built around vLLM, where he is CEO.
SGLang treats LLM applications as structured programs. A frontend language expresses multiple generation calls, control flow and constrained output. It is co-designed with a runtime whose RadixAttention reuses the key-value cache across calls and whose compressed finite-state machines speed up structured decoding. SGLang has become a major, widely adopted open-source inference engine used to serve models in production.
People: Created by Lianmin Zheng, my PhD student, together with Ying Sheng, with collaborators including Shiyi Cao, a current student in the group.
Alpa automatically combined inter-operator (pipeline) and intra-operator (tensor) parallelism, compiling large models onto accelerator clusters without hand-written parallelization plans. AlpaServe (OSDI 2023) showed how model parallelism can also improve serving through statistical multiplexing.
People: Led by my student Lianmin Zheng with Zhuohan Li, Hao Zhang and collaborators. Zhuohan Li led AlpaServe.
Anonymous, crowdsourced side-by-side comparisons of models on real user prompts, aggregated into statistically grounded ratings: an open, continually updated benchmark. Companion work introduced MT-Bench and LLM-as-a-Judge, using strong models to evaluate other models, and released LMSYS-Chat-1M, a dataset of real-world conversations.
People: I started Chatbot Arena within LMSYS and ran its early experiments together with students and collaborators, including Lianmin Zheng, Wei-Lin Chiang and Ying Sheng. My student Dacheng Li also contributed.
Gorilla trained LLMs, with retrieval over documentation, to write correct API calls, reducing hallucinated calls. The Berkeley Function Calling Leaderboard (BFCL) became a widely used benchmark for tool use, and RAFT studied how to adapt models to retrieval over specific domains.
If many hard problems in agents are solved, systems, human oversight and learning itself will have to change.
Today's bet is that many of the hard problems in AI agents will be solved, and that highly capable agents will arrive in very large numbers. That raises questions beyond building any single agent:
Systems for thousands of capable agents. Systems have been designed around human users and conventional programs. How should interfaces and abstractions change for this new, more capable class of user, and how will systems handle thousands of agents at once?
Humans steering teams of agents. How can people direct vast teams of agents that are faster and more sophisticated than the people leading them? We may have to let go of detailed understanding, and perhaps some basic design principles, in favor of higher-level abstractions that have yet to be developed.
Negotiation and consensus. When everyone has capable agents, how will people with competing goals negotiate and find consensus?
Learning from teachers. Beyond learning from data, how can models exchange knowledge and capabilities? What does it mean to learn how to teach, and how efficient can knowledge transfer become?
Current projects are early steps on these questions: data systems redesigned for agents as their primary users, serving engines for multi-turn agents, studies of how agents cooperate and compete, and small advisor models that learn to steer larger ones. The group also continues to work on post-training and on new neural architectures. That work returns to its oldest question: how to co-design models with the systems that run them.
AI agents
Systems for agents that keep memory, call tools, run for many steps and interact with each other.
MemGPT
LLMs as operating systems: persistent memory for agents
MemGPT's virtual context management pages information between a model's limited context window and external storage, much as an operating system manages memory. This lets agents maintain persistent memory across long conversations and documents. Follow-on work on sleep-time compute lets agents reason about their context between interactions.
People: Created by my students Charles Packer, Sarah Wooders and Kevin Lin with collaborators. Charles and Sarah co-founded Letta to continue the work, and Kevin joined them. Kevin led the sleep-time compute study.
Scheduling and caching for multi-turn, tool-using agents
Agents interleave model calls with tool calls, which breaks the assumptions of request-level LLM serving. Autellix schedules agent programs as a whole. Continuum keeps a multi-turn agent's key-value cache alive across tool calls. vCache offers verified semantic prompt caching.
Measuring agents in the wild and diagnosing failures
Search Arena extends human-preference evaluation to search-augmented LLMs. MAST is a taxonomy of why multi-agent LLM systems fail. Characterizing Agents in Production studies how agents are actually built and deployed.
People: Search Arena was led by Mihran Miroyan with Patrick Wu. MAST and the production study were led by collaborators, with contributions from students including Huanzhi Mao and Shu Liu.
Cooperate to Compete studies strategic coordination among LLM agents in a multi-agent conquest game, part of the group's broader work on how AI systems interact with one another.
Retrieval, memory and data systems designed for agents
If agents become the primary users of data systems, those systems should be redesigned around them. LEANN is a storage-efficient vector index for retrieval on personal devices.
People: The agent-first data systems vision was led by Shu Liu with Alan Zhu and collaborators. LEANN was led by Yichuan Wang.
How models keep improving after pre-training: reinforcement learning, steering by smaller models, continual learning and transferring knowledge between models.
SkyRL
Reinforcement learning for long-horizon agents
An open framework for training LLM agents with reinforcement learning on multi-turn, tool-using tasks such as software engineering. SkyRL-Agent focuses on efficient training for multi-turn, long-horizon agents.
People: Initially co-developed by students in the group and collaborators in the Sky Computing Lab. SkyRL-Agent was led by Shiyi Cao and Dacheng Li.
Steering, adapting and continually improving models
Advisor Models trains small models that steer black-box LLMs toward specific users and tasks. Continual Learning Bench measures whether agents actually improve from experience over time.
Post-training tends to narrow the range of outputs a model will produce. SimpleStrat diversifies generation by stratifying the space of possible answers and sampling across strata. It is part of a broader study of how post-training affects generative diversity.
People: Led by Justin Wong, whose thesis studied how to mitigate the effects of post-training on generative diversity.
New model architectures, attention mechanisms and kernels for efficient, scalable models.
MRNN
Non-linear RNNs with matrix-valued states
A recurrent architecture that keeps a non-linear recurrence but makes the hidden state a matrix rather than a vector, aimed at scalable language modeling as an alternative to attention-only designs.
HashAttention exploits semantic sparsity to find the few tokens that matter for each query. vAttention adds statistical guarantees to sparse attention by sampling. SLA combines sparse and linear attention in diffusion transformers and underlies TurboDiffusion, which accelerates video generation by 100–200×.
People: HashAttention and vAttention were led by postdoc Aditya Desai. SLA and TurboDiffusion were led by Jintao Zhang and collaborators at Tsinghua, including former postdoc Jianfei Chen.
K-Search uses LLMs to generate high-performance GPU kernels, guiding the search with an intrinsic world model that co-evolves with the candidate kernels.
Research questions often move through the group with the students who pursue them. These are a few of the paths former students followed at Berkeley and beyond. The work in each case was theirs, done with many collaborators.
From graph processing to the first generation of prediction-serving systems. His thesis, The Design and Implementation of Low-Latency Prediction Serving Systems, laid the foundation for the serving line of work.
Carried prediction serving from Clipper into Ray Serve and then into open-source LLM inference with vLLM. His thesis is Building Open Source Inference Serving Systems.
From automatic parallelism for large models, to LLM serving and human-preference evaluation, to creating SGLang. His thesis is Scalable and Efficient Systems for Large Deep Learning Models.
From treating LLMs as operating systems to building persistent, stateful agents. His thesis is Building Agentic Systems in an Era of Large Language Models.
A Research Community
These projects share more than a line of ideas; they share people. Most were led by PhD students I advised, working with collaborators across Berkeley's AMPLab, RISELab, Sky Computing Lab and LMSYS, and at other institutions. Former students have gone on to faculty positions, AI research labs and companies they founded, and current students are carrying the work into AI agents.