Skip to content

Repository files navigation

DOE_HPC_Bootcamp_2026

Welcome to the repository for the DOE High-Performance Computing (HPC) Bootcamp 2026. This project is designed to help participants build practical HPC workflow skills, including running GPU simulations, validating outputs, tracking reproducibility, and learning the basics of performance analysis on Perlmutter.

This repository contains the sim2spec project, developed as part of the DOE HPC Bootcamp at Argonne National Laboratory, from August 9-14 in St. Charles, IL. For broader context on the bootcamp, see the Argonne Introduction to HPC Bootcamp page.

New to HPC or need a refresher before the bootcamp? Check out the Bootcamp Prep Pack — a curated set of prerequisite materials covering the Linux command line, Python basics, and HPC concepts to help you get the most out of the week.

sim2spec

sim2spec is a lightweight workflow wrapper around DUNE/larnd-sim for running, validating, comparing, and profiling GPU-based detector simulations on Perlmutter.

This repository is designed for students who are new to HPC workflows and want a structured way to:

  • install and run a GPU simulation workflow,
  • generate QA summaries and plots,
  • compare multiple run configurations,
  • track provenance and reproducibility,
  • and profile performance on Perlmutter (A100 GPU).

This project does not replace larnd-sim. Instead, it organizes the workflow around it.

How to use this project

This repository is organized for HPC beginners. Start by reading the concept sections below, then open the daily project files as exercises. The goal is not to memorize every command; the goal is to understand what each command is asking the HPC system to do and how to check whether it worked.

flowchart TD
    subgraph row1[" "]
        direction LR
        A[Your laptop or browser] --> B[NERSC login node]
        B --> C[Slurm scheduler]
        C --> D[Perlmutter GPU compute node]
    end

    subgraph row2[" "]
        direction LR
        E[larnd-sim run] --> F[output.h5]
        F --> G[QA metrics and plots]
        G --> H[Sweep comparison]
        H --> I[Profiling and final summary]
    end

    row1 --> row2
Loading

Beginner concepts map

HPC system

  • HPC means high-performance computing: using shared supercomputing resources for work that is too large, slow, or specialized for a laptop.
  • Perlmutter is the NERSC supercomputer used by this project.
  • Login node is where you log in, edit files, submit jobs, and manage the project. Do not run heavy GPU simulations directly on login nodes.
  • Compute node is where Slurm runs your actual job. GPU simulations should run on GPU compute nodes.
  • $PSCRATCH is a high-performance scratch filesystem for active work and output files.

Scheduling

  • Slurm is the workload manager. It decides when and where your job runs.
  • salloc requests an interactive allocation. Use it when you want a live shell on a compute node.
  • sbatch submits a batch job script. Use it when you want the system to run the workflow without keeping an interactive shell open.
  • srun launches work inside an allocation or asks Slurm to run one command on compute resources.

For a full reference of sbatch flags used in this project, see scripts/sbash_flags.md.

Software environment

  • Environment modules load site-provided software such as Python.
  • Python virtual environment keeps this project's Python packages separate from other projects.
  • CUDA is NVIDIA's GPU programming platform.
  • CuPy provides NumPy-like arrays that run on NVIDIA GPUs.
  • Numba is a Python JIT compiler often used in GPU and performance-oriented Python workflows.

Simulation workflow

  • larnd-sim is the detector simulation package that performs the core GPU-based simulation.
  • sim2spec is the wrapper in this repository. It organizes larnd-sim runs, QA, sweeps, provenance, and profiling into a beginner-friendly workflow.
  • HDF5 is the file format used for large structured simulation input and output files.
  • output.h5 is the main simulation result file produced by a run.

Validation and reproducibility

  • QA means quality assurance: quick checks that the output exists and contains reasonable datasets, counts, ranges, and plots.
  • Validation plots help connect numbers in QA metrics to physical behavior such as charge timing, event activity, and light waveforms.
  • Manifest means a machine-readable record of how a run was produced.
  • Provenance means the information needed to understand and reproduce a result: input file, code version, seed, environment, command, and output path.
  • Random seed controls stochastic parts of a simulation so different variants can be compared systematically.

Performance

  • Profiling measures where time is spent.
  • Nsight Systems is NVIDIA's timeline profiler for CPU/GPU applications.
  • Kernel means a function launched on the GPU.
  • Wall time is the elapsed time you wait for a run to finish.

Exercise guides

  • Project_1_ReadMe.md — Day 1: environment setup, install, and smoke test
  • Project_2_ReadMe.md — Day 2: baseline run, QA, and validation plots
  • Project_3_ReadMe.md — Day 3: parameter sweeps, provenance tracking, and first profiling with Nsight Systems
  • Project_4_ReadMe.md — Day 4: one measurable improvement (TPB) and kernel profiling with Nsight Compute
  • Project_5_ReadMe.md — Day 5: final cross-day comparison and summary
  • sim2spec_perlmutter_bootcamp.ipynb — interactive notebook for participants who prefer to complete the exercises in Jupyter instead of the terminal. On Perlmutter, use the available Python environment and select NERSC Python as the notebook kernel. It is meant for the exercises only. Please still read daily Project_N_ReadMe.md for more HPC background and context.
  • Extended or optional: the scripts/ folder includes batch scripts for NERSC users who want to submit jobs directly with sbatch, as well as Python helper scripts used in the daily exercises.

Quick start

1. Clone the repository

export MYWORKDIR=$PSCRATCH/HPC_intro
mkdir -p "$MYWORKDIR"
cd "$MYWORKDIR"

git clone https://github.com/madantimalsina/sim2spec.git
cd sim2spec

2. Set up the environment

source setup.sh
bash install.sh

3. Activate the virtual environment

source setup.sh
source "$venv_name/bin/activate"

4. Quick environment check

Run these checks before trying a real workflow step:

python -c "import fire; print('fire ok')"
python -c "import cupy as cp; print(int(cp.arange(10).sum()))"
python -c "import larndsim; print('larndsim ok')"

If these work, your Python environment is in good shape.

Before you begin

You will need a NERSC compute allocation and access to Perlmutter. For GPU work, request an interactive node or submit a batch job before running simulations.

For short setup checks and baseline tests:

salloc -C gpu -q interactive -t 00:30:00 -A <your_account> --gpus=1 --ntasks=1 --cpus-per-task=8

For example:

salloc -C gpu -q interactive -t 00:60:00 -A m4388 --gpus=1 --ntasks=1 --cpus-per-task=8

Basic workflow examples

Note: follow the matching daily guide for the block you are working on.

Baseline run

export WORKDIR=$PWD
export LARNDSIM_DIR=$WORKDIR/larnd-sim
export INPUT_H5=$WORKDIR/input/MiniRun5_1E19_RHC.convert2h5.0000123.EDEPSIM.hdf5
export HDF5_USE_FILE_LOCKING=0
export LARNDSIM_DISABLE_CUPY_MEMPOOL=1
export OUTBASE=$WORKDIR/runs

mkdir -p "$OUTBASE"

sim2spec run \
  --larndsim-dir "$LARNDSIM_DIR" \
  --config 2x2 \
  --input "$INPUT_H5" \
  --outdir "$OUTBASE/day2_baseline" \
  --n-events 5

QA on an existing run

sim2spec qa --run-dir "$OUTBASE/day2_baseline/run"

Sweep run

sim2spec sweep \
  --larndsim-dir "$LARNDSIM_DIR" \
  --config 2x2 \
  --input "$INPUT_H5" \
  --outdir "$OUTBASE/day3_sweep" \
  --sweep "$WORKDIR/configs/sweep.yaml" \
  --n-events 3

Profiling run (Day 3 — TPB = 4 baseline)

export LARNDSIM_DISABLE_CUPY_MEMPOOL=1

sim2spec run \
  --larndsim-dir "$LARNDSIM_DIR" \
  --config 2x2 \
  --input "$INPUT_H5" \
  --outdir "$OUTBASE/day3_profile_baseline" \
  --n-events 5 \
  --profiler nsys

Write profile summary

sim2spec profile --run-dir "$OUTBASE/day3_profile_baseline/run"

Profiling run (Day 4 — TPB = 64 comparison)

sim2spec run \
  --larndsim-dir "$LARNDSIM_DIR" \
  --config 2x2 \
  --input "$INPUT_H5" \
  --outdir "$OUTBASE/day4_profile_tpb64" \
  --n-events 5 \
  --profiler nsys

Kernel profiling with Nsight Compute (Day 4)

sim2spec run \
  --larndsim-dir "$LARNDSIM_DIR" \
  --config 2x2 \
  --input "$INPUT_H5" \
  --outdir "$OUTBASE/day4_ncu" \
  --n-events 3 \
  --profiler ncu

Prefer batch jobs (extended or optional)?

Learning how to submit sbatch jobs is important for HPC users, but do not worry about this for now. You can look at this section later.

Every GPU step also has a corresponding sbatch script in scripts/. For example:

sbatch scripts/sbatch_day1_smoke.sh
sbatch scripts/sbatch_day2_baseline.sh
sbatch scripts/sbatch_day3_sweep.sh
sbatch scripts/sbatch_day3_profile_baseline.sh
sbatch scripts/sbatch_day4_profile_compare.sh
sbatch scripts/sbatch_day4_ncu.sh

Remember to replace <your_account> with your NERSC project account before submitting.

Repository structure

Below is a beginner-friendly explanation of the main files and folders in the repository.

Very small file.
It mainly defines the package version and marks src/ as the Python package location.

Why it matters:

  • helps Python treat the source code as an installable package
  • provides a clean package entry point

This is the main command-line entry point.

It defines the commands:

  • sim2spec run
  • sim2spec sweep
  • sim2spec qa
  • sim2spec profile

What each command does:

run

  • runs a single larnd-sim job
  • writes output to a run directory
  • saves a manifest and command file

sweep

  • loads multiple variants from a YAML file
  • runs one simulation per variant
  • automatically runs QA for each

qa

  • reads one run directory
  • generates metrics and plots

profile

  • parses profiling outputs for one run

Why it matters:

  • this file is the public interface of the project
  • when a user types sim2spec ..., this is what executes

This file is responsible for actually launching larnd-sim.

What it does:

  • builds the command used to call larnd-sim
  • creates the run directory
  • saves the command and manifest
  • optionally wraps the run with profiling tools such as nsys

Why it matters:

  • this is the main bridge between sim2spec and larnd-sim
  • if you want to understand how a run is executed, start here

This is the quality-assurance module.

What it does:

  • opens the output HDF5 file
  • checks for expected datasets
  • computes summary metrics such as:
    • packet counts
    • ADC statistics
    • timestamp range
    • light waveform counts if present
  • creates quick plots for validation

Why it matters:

  • this is the first layer of output validation
  • it helps answer the question: "does this run look reasonable?"

This module handles reproducibility metadata.

What it does:

  • records timestamps and environment information
  • records selected environment variables
  • collects git information for the larnd-sim checkout
  • helps build manifest.json for each run

Why it matters:

  • reproducibility is a major part of the workflow
  • this file makes each run easier to understand and reproduce later

This is the profiling helper module.

What it does:

  • finds profiling outputs, especially nsys results
  • runs summary commands such as nsys stats
  • saves a simplified JSON report

Why it matters:

  • raw profiling outputs can be hard to read directly
  • this module makes them easier to compare between runs

This file handles sweep-related configuration logic.

What it does:

  • reads YAML files
  • loads sweep variants from configs/sweep.yaml
  • writes updated YAML if needed

Why it matters:

  • the sweep workflow depends on a clean way to define multiple run variants

This file contains small helper utilities used across the project.

Typical examples include:

  • creating directories safely
  • reading and writing JSON
  • collecting timestamps
  • merging dictionaries

Why it matters:

  • it keeps repeated helper logic out of the main workflow files

This file defines the parameter sweep used by sim2spec sweep.

What it does:

  • lists four named variants, each with a distinct random seed
  • the seed drives stochastic variation so each variant produces observably different outputs

Why it matters:

  • this is how multiple runs are compared in a controlled, reproducible way

This folder stores input files used for the workflow.

For this project, it is expected to contain:

  • MiniRun5_1E19_RHC.convert2h5.0000123.EDEPSIM.hdf5

Why it matters:

  • keeping the input in a predictable place makes the notebook and shell scripts easier to follow

This folder contains two types of scripts:

  • Batch scripts (sbatch_day*.sh) — submit GPU jobs to Slurm directly with sbatch
  • Python helper scripts — called from the daily exercises to compare metrics, extract provenance, and save CSV outputs:
    • save_metrics_csv.py — Day 2: save QA metrics to CSV
    • compare_sweep_metrics.py — Day 3: compare packet counts and ADC stats across sweep variants
    • extract_sweep_provenance.py — Day 3: print seed, config, and git commit from each manifest
    • save_sweep_comparison_csv.py — Day 3: save full sweep comparison table to comparison.csv
    • compare_profile_runs.py — Day 4: compare wall time and output file size between profile runs
    • compare_baseline_vs_sweep.py — Day 5: compare baseline and sweep metrics side by side

Why it matters:

  • keeps inline Python out of the README and Jupyter cells
  • makes each step runnable with a single python scripts/<name>.py call

This is the lightweight environment setup script.

What it does:

  • unloads any conflicting Python module
  • loads python/3.11
  • defines the virtual-environment name
  • prepares shell variables used by install.sh

Why it matters:

  • this is the recommended first command to source when starting a terminal session for the project

This is the main installation script for terminal users.

What it does:

  • creates a virtual environment
  • installs Python dependencies
  • installs sim2spec
  • clones and installs larnd-sim (skips clone if already present)

Why it matters:

  • this is the easiest way for terminal users to get started without following the full notebook

Standalone script for producing validation plots directly from an output HDF5 file.

python plot_validation.py "$OUTBASE/day2_baseline/run/output.h5" --outdir "$OUTBASE/day2_baseline/run/validation_plots"

Why it matters:

  • a quick way to visualise any run output without running the full QA pipeline

This file is the top-level overview of the repository.

Why it matters:

  • it gives a quick map of the project
  • it points students to the detailed guides and notebooks

These are the focused day-by-day project guides for terminal users.

What they contain:

  • Day 1: environment setup, install, and smoke test
  • Day 2: baseline run, QA, and validation plots
  • Day 3: parameter sweeps, provenance tracking, and first profiling with Nsight Systems
  • Day 4: one measurable improvement (TPB) and kernel profiling with Nsight Compute
  • Day 5: final cross-day comparison and summary

Why they matter:

  • each file keeps one bootcamp work block short, focused, and easier to follow during project time

The main student-facing notebook. Mirrors the Project_1_ReadMe.md through Project_5_ReadMe.md day guides with executable cells and srun-based GPU dispatch so students can run everything from inside JupyterHub.

Why it matters:

  • the recommended path for students who prefer notebooks over the terminal
  • keeps the notebook path shorter than the readmes while producing the same core outputs

Suggested reading order

If you are new to the repository, a good order is:

  1. README.md
  2. Project_1_ReadMe.md through Project_5_ReadMe.md
  3. sim2spec_perlmutter_bootcamp.ipynb — if you prefer notebooks
  4. src/cli.py
  5. src/runner.py
  6. src/qa.py
  7. src/provenance.py
  8. configs/sweep.yaml

References

Bootcamp and computing resources:

Simulation and detector workflow resources:

Tools used by this workflow:

Notes

  • This wrapper does not replace larnd-sim; it organizes and validates runs around it.
  • The project is designed for Perlmutter-style Python 3.11 + venv usage.
  • For Jupyter on Perlmutter, we will use the available Python environment and corresponding Jupyter kernel. Select NERSC Python as the notebook kernel.

About

Lightweight workflow wrapper around DUNE/larnd-sim for GPU detector simulation on Perlmutter — 5-day HPC bootcamp project

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages