SPLASH 2026
Sun 4 - Fri 9 October 2026 Oakland, California, United States
co-located with SPLASH/ISSTA 2026

This program is tentative and subject to change.

Tue 6 Oct 2026 16:24 - 16:42 at Junior Ballroom 1&2 - Runtime Systems and Performance Chair(s): Rohan Padhye

Accelerator design languages (ADLs), high-level languages that compile to hardware units, help domain experts quickly design efficient application-specific hardware. ADL compilers optimize datapaths and convert software-like control flow constructs into control paths. Such compilers are necessarily complex and often unpredictable: they must bridge the wide semantic gap between high-level semantics and cycle-level schedules, and they typically rely on advanced heuristics to optimize circuits. The resulting performance can be difficult to control, requiring guesswork to find and resolve performance problems in the generated hardware. We conjecture that ADL compilers will never be perfect: some performance unpredictability is endemic to the problem they solve.

In lieu of compiler perfection, we argue for \emph{compiler understanding tools} that give ADL programmers insight into how the compiler's decisions affect performance. We introduce Petal, a cycle-level profiler for ADLs that compile to the Calyx intermediate language (IL). Petal instruments the Calyx code with probes and then analyzes the trace from a register-transfer-level simulation. It then maps the events in the trace back to high-level control constructs in the Calyx code to determine when each construct was active. Petal processes that information into a trace of \emph{call trees}, each representing active events in a specific cycle and their relationships. Lastly, Petal uses metadata generated by the ADL compiler to construct an ADL-level profile. Using case studies, we demonstrate that Petal's cycle-level profiles can identify performance problems in existing accelerator designs. We show that these insights can also guide developers toward optimizations that the compiler was unable to perform automatically, including a reduction by 46.9% of total cycles for one application.

This program is tentative and subject to change.

Tue 6 Oct

Displayed time zone: Pacific Time (US & Canada) change

15:30 - 17:00
Runtime Systems and PerformanceOOPSLA at Junior Ballroom 1&2
Chair(s): Rohan Padhye Carnegie Mellon University and Antithesis
15:30
18m
Talk
Uncovering Hidden Memory Costs for Garbage Collection
OOPSLA
Sudhanshu Agarwal University of Illinois at Urbana-Champaign, Saugata Ghose University of Illinois at Urbana-Champaign
DOI
15:48
18m
Talk
Designing GPU Data Structures for Efficient Memory Oversubscription
OOPSLA
Vipin Patel IIT Kanpur, Srinjoy Sarkar IIT Kanpur, Swarnendu Biswas IIT Kanpur, Mainak Chaudhuri IIT Kanpur
DOI
16:06
18m
Talk
Bonsai: Efficient and Optimal Automatic Tensor Rematerialization for Memory-Constrained DNN Training
OOPSLA
Dat Nguyen Texas A&M University, Vasudha Devarakonda Texas A&M University, Anxiao Jiang Texas A&M University, Khanh Nguyen Texas A&M University
DOI
16:24
18m
Talk
Understanding Accelerator Compilers via Performance Profiling
OOPSLA
Ayaka Yorihiro Cornell University, Griffin Berlstein Cornell University, Pedro Pontes García Cornell University, Kevin Laeufer Cornell University, Adrian Sampson Cornell University
DOI
16:42
18m
Talk
A Language Approach to Fine-Grained Microarchitectural Observation
OOPSLA
DOI
Hide past events