What is it about?
Performance anomalies cause significant variations in job execution time across large-scale computing systems. Even though their impact is significant (e.g., up to a 7x increase in execution time), anomalies infrequently occur, which makes it harder to collect data about their characteristics. This paper addresses this lack of a common method for creating relevant performance anomalies by introducing HPAS, an HPC Performance Anomaly Suite, consisting of anomaly generators for the major subsystems in HPC systems.
Featured Image
Why is it important?
Anomalies cause performance variations, and these variations can cause up to a 7x increase in the execution time of a job within the same system.
Read the Original
This page is a summary of: HPAS, August 2019, ACM (Association for Computing Machinery),
DOI: 10.1145/3337821.3337907.
You can read the full text:
Resources
Github
This repository holds the anomaly suite presented in the paper "HPAS: An HPC Performance Anomaly Suite for Reproducing Performance Variations"[1]. It consists of a set of synthetic anomalies that reproduce common root causes of performance variations in supercomputers: CPU contention Cache evictions Memory bandwidth interference Memory intensive processes Memory leaks Network contention I/O metadata server contention I/O storage server contention
Paper
Paper URL
Contributors
The following have contributed to this page







