SPLASH 2026
Sun 4 - Fri 9 October 2026 Oakland, California, United States
co-located with SPLASH/ISSTA 2026

This program is tentative and subject to change.

Tue 6 Oct 2026 16:06 - 16:24 at Junior Ballroom 1&2 - Runtime Systems and Performance Chair(s): Rohan Padhye

GPU memory is increasingly the primary bottleneck in scaling deep neural network (DNN) training, where the activation tensors footprint of a model may exceed the memory capacity. Tensor recomputation is a powerful technique that trades additional computation for reduced peak memory usage. However, existing approaches face a fundamental tension between performance optimality and computational scalability. On the one hand, solvers leverage Integer Linear Programming (ILP) to provide mathematically optimal solutions but suffer from the combinatorial explosion of the search space and thus become intractable for modern DNN models. On the other hand, heuristics-based approaches achieve scalability but sacrifice optimality altogether, resulting in suboptimal execution schedules.

The root cause of these inefficiencies in the state of the art is the mismatch in abstraction. This paper introduces Bonsai, a framework that tackles this scalability-granularity tension. At the heart of Bonsai is a novel abstraction of operator segmentation that breaks the computation graph into flexible, variable-sized units to enable a lightweight yet effective segment-based ILP formulation. By having segments, Bonsai collapses the search space and prunes redundant solutions that stall existing solvers. This abstraction enables Bonsai to maintain a holistic view of the entire model, ensuring that no optimization opportunity is lost while reducing the number of decision variables by orders of magnitude. The evaluation across a diverse set of DNN architectures and models demonstrates that Bonsai scales to real-world models, is up to 10.13$\times$ lower solver cost than state-of-the-art ILP solvers, and delivers up to 11% improvement in training throughput.

This program is tentative and subject to change.

Tue 6 Oct

Displayed time zone: Pacific Time (US & Canada) change

15:30 - 17:00
Runtime Systems and PerformanceOOPSLA at Junior Ballroom 1&2
Chair(s): Rohan Padhye Carnegie Mellon University and Antithesis
15:30
18m
Talk
Uncovering Hidden Memory Costs for Garbage Collection
OOPSLA
Sudhanshu Agarwal University of Illinois at Urbana-Champaign, Saugata Ghose University of Illinois at Urbana-Champaign
DOI
15:48
18m
Talk
Designing GPU Data Structures for Efficient Memory Oversubscription
OOPSLA
Vipin Patel IIT Kanpur, Srinjoy Sarkar IIT Kanpur, Swarnendu Biswas IIT Kanpur, Mainak Chaudhuri IIT Kanpur
DOI
16:06
18m
Talk
Bonsai: Efficient and Optimal Automatic Tensor Rematerialization for Memory-Constrained DNN Training
OOPSLA
Dat Nguyen Texas A&M University, Vasudha Devarakonda Texas A&M University, Anxiao Jiang Texas A&M University, Khanh Nguyen Texas A&M University
DOI
16:24
18m
Talk
Understanding Accelerator Compilers via Performance Profiling
OOPSLA
Ayaka Yorihiro Cornell University, Griffin Berlstein Cornell University, Pedro Pontes García Cornell University, Kevin Laeufer Cornell University, Adrian Sampson Cornell University
DOI
16:42
18m
Talk
A Language Approach to Fine-Grained Microarchitectural Observation
OOPSLA
DOI
Hide past events