Welcome to the repository for the DOE High-Performance Computing (HPC) Bootcamp 2026. This project is designed to help participants build practical HPC workflow skills, including running GPU simulations, validating outputs, tracking reproducibility, and learning the basics of performance analysis on Perlmutter.
This repository contains the sim2spec project, developed as part of the DOE HPC Bootcamp at Argonne National Laboratory, from August 9-14 in St. Charles, IL. For broader context on the bootcamp, see the Argonne Introduction to HPC Bootcamp page.
New to HPC or need a refresher before the bootcamp? Check out the Bootcamp Prep Pack — a curated set of prerequisite materials covering the Linux command line, Python basics, and HPC concepts to help you get the most out of the week.
sim2spec is a lightweight workflow wrapper around DUNE/larnd-sim for running, validating, comparing, and profiling GPU-based detector simulations on Perlmutter.
This repository is designed for students who are new to HPC workflows and want a structured way to:
- install and run a GPU simulation workflow,
- generate QA summaries and plots,
- compare multiple run configurations,
- track provenance and reproducibility,
- and profile performance on Perlmutter (A100 GPU).
This project does not replace larnd-sim. Instead, it organizes the workflow around it.
This repository is organized for HPC beginners. Start by reading the concept sections below, then open the daily project files as exercises. The goal is not to memorize every command; the goal is to understand what each command is asking the HPC system to do and how to check whether it worked.
flowchart TD
subgraph row1[" "]
direction LR
A[Your laptop or browser] --> B[NERSC login node]
B --> C[Slurm scheduler]
C --> D[Perlmutter GPU compute node]
end
subgraph row2[" "]
direction LR
E[larnd-sim run] --> F[output.h5]
F --> G[QA metrics and plots]
G --> H[Sweep comparison]
H --> I[Profiling and final summary]
end
row1 --> row2
- HPC means high-performance computing: using shared supercomputing resources for work that is too large, slow, or specialized for a laptop.
- Perlmutter is the NERSC supercomputer used by this project.
- Login node is where you log in, edit files, submit jobs, and manage the project. Do not run heavy GPU simulations directly on login nodes.
- Compute node is where Slurm runs your actual job. GPU simulations should run on GPU compute nodes.
$PSCRATCHis a high-performance scratch filesystem for active work and output files.
- Slurm is the workload manager. It decides when and where your job runs.
sallocrequests an interactive allocation. Use it when you want a live shell on a compute node.sbatchsubmits a batch job script. Use it when you want the system to run the workflow without keeping an interactive shell open.srunlaunches work inside an allocation or asks Slurm to run one command on compute resources.
For a full reference of sbatch flags used in this project, see scripts/sbash_flags.md.
- Environment modules load site-provided software such as Python.
- Python virtual environment keeps this project's Python packages separate from other projects.
- CUDA is NVIDIA's GPU programming platform.
- CuPy provides NumPy-like arrays that run on NVIDIA GPUs.
- Numba is a Python JIT compiler often used in GPU and performance-oriented Python workflows.
larnd-simis the detector simulation package that performs the core GPU-based simulation.sim2specis the wrapper in this repository. It organizeslarnd-simruns, QA, sweeps, provenance, and profiling into a beginner-friendly workflow.- HDF5 is the file format used for large structured simulation input and output files.
output.h5is the main simulation result file produced by a run.
- QA means quality assurance: quick checks that the output exists and contains reasonable datasets, counts, ranges, and plots.
- Validation plots help connect numbers in QA metrics to physical behavior such as charge timing, event activity, and light waveforms.
- Manifest means a machine-readable record of how a run was produced.
- Provenance means the information needed to understand and reproduce a result: input file, code version, seed, environment, command, and output path.
- Random seed controls stochastic parts of a simulation so different variants can be compared systematically.
- Profiling measures where time is spent.
- Nsight Systems is NVIDIA's timeline profiler for CPU/GPU applications.
- Kernel means a function launched on the GPU.
- Wall time is the elapsed time you wait for a run to finish.
- Project_1_ReadMe.md — Day 1: environment setup, install, and smoke test
- Project_2_ReadMe.md — Day 2: baseline run, QA, and validation plots
- Project_3_ReadMe.md — Day 3: parameter sweeps, provenance tracking, and first profiling with Nsight Systems
- Project_4_ReadMe.md — Day 4: one measurable improvement (TPB) and kernel profiling with Nsight Compute
- Project_5_ReadMe.md — Day 5: final cross-day comparison and summary
- sim2spec_perlmutter_bootcamp.ipynb — interactive notebook for participants who prefer to complete the exercises in Jupyter instead of the terminal. On Perlmutter, use the available Python environment and select
NERSC Pythonas the notebook kernel. It is meant for the exercises only. Please still read dailyProject_N_ReadMe.mdfor more HPC background and context. - Extended or optional: the
scripts/folder includes batch scripts for NERSC users who want to submit jobs directly withsbatch, as well as Python helper scripts used in the daily exercises.
export MYWORKDIR=$PSCRATCH/HPC_intro
mkdir -p "$MYWORKDIR"
cd "$MYWORKDIR"
git clone https://github.com/madantimalsina/sim2spec.git
cd sim2specsource setup.sh
bash install.shsource setup.sh
source "$venv_name/bin/activate"Run these checks before trying a real workflow step:
python -c "import fire; print('fire ok')"
python -c "import cupy as cp; print(int(cp.arange(10).sum()))"
python -c "import larndsim; print('larndsim ok')"If these work, your Python environment is in good shape.
You will need a NERSC compute allocation and access to Perlmutter. For GPU work, request an interactive node or submit a batch job before running simulations.
For short setup checks and baseline tests:
salloc -C gpu -q interactive -t 00:30:00 -A <your_account> --gpus=1 --ntasks=1 --cpus-per-task=8For example:
salloc -C gpu -q interactive -t 00:60:00 -A m4388 --gpus=1 --ntasks=1 --cpus-per-task=8Note: follow the matching daily guide for the block you are working on.
export WORKDIR=$PWD
export LARNDSIM_DIR=$WORKDIR/larnd-sim
export INPUT_H5=$WORKDIR/input/MiniRun5_1E19_RHC.convert2h5.0000123.EDEPSIM.hdf5
export HDF5_USE_FILE_LOCKING=0
export LARNDSIM_DISABLE_CUPY_MEMPOOL=1
export OUTBASE=$WORKDIR/runs
mkdir -p "$OUTBASE"
sim2spec run \
--larndsim-dir "$LARNDSIM_DIR" \
--config 2x2 \
--input "$INPUT_H5" \
--outdir "$OUTBASE/day2_baseline" \
--n-events 5sim2spec qa --run-dir "$OUTBASE/day2_baseline/run"sim2spec sweep \
--larndsim-dir "$LARNDSIM_DIR" \
--config 2x2 \
--input "$INPUT_H5" \
--outdir "$OUTBASE/day3_sweep" \
--sweep "$WORKDIR/configs/sweep.yaml" \
--n-events 3export LARNDSIM_DISABLE_CUPY_MEMPOOL=1
sim2spec run \
--larndsim-dir "$LARNDSIM_DIR" \
--config 2x2 \
--input "$INPUT_H5" \
--outdir "$OUTBASE/day3_profile_baseline" \
--n-events 5 \
--profiler nsyssim2spec profile --run-dir "$OUTBASE/day3_profile_baseline/run"sim2spec run \
--larndsim-dir "$LARNDSIM_DIR" \
--config 2x2 \
--input "$INPUT_H5" \
--outdir "$OUTBASE/day4_profile_tpb64" \
--n-events 5 \
--profiler nsyssim2spec run \
--larndsim-dir "$LARNDSIM_DIR" \
--config 2x2 \
--input "$INPUT_H5" \
--outdir "$OUTBASE/day4_ncu" \
--n-events 3 \
--profiler ncuLearning how to submit sbatch jobs is important for HPC users, but do not worry about this for now. You can look at this section later.
Every GPU step also has a corresponding sbatch script in scripts/. For example:
sbatch scripts/sbatch_day1_smoke.sh
sbatch scripts/sbatch_day2_baseline.sh
sbatch scripts/sbatch_day3_sweep.sh
sbatch scripts/sbatch_day3_profile_baseline.sh
sbatch scripts/sbatch_day4_profile_compare.sh
sbatch scripts/sbatch_day4_ncu.shRemember to replace <your_account> with your NERSC project account before submitting.
Below is a beginner-friendly explanation of the main files and folders in the repository.
Very small file.
It mainly defines the package version and marks src/ as the Python package location.
Why it matters:
- helps Python treat the source code as an installable package
- provides a clean package entry point
This is the main command-line entry point.
It defines the commands:
sim2spec runsim2spec sweepsim2spec qasim2spec profile
What each command does:
run
- runs a single
larnd-simjob - writes output to a run directory
- saves a manifest and command file
sweep
- loads multiple variants from a YAML file
- runs one simulation per variant
- automatically runs QA for each
qa
- reads one run directory
- generates metrics and plots
profile
- parses profiling outputs for one run
Why it matters:
- this file is the public interface of the project
- when a user types
sim2spec ..., this is what executes
This file is responsible for actually launching larnd-sim.
What it does:
- builds the command used to call
larnd-sim - creates the run directory
- saves the command and manifest
- optionally wraps the run with profiling tools such as
nsys
Why it matters:
- this is the main bridge between
sim2specandlarnd-sim - if you want to understand how a run is executed, start here
This is the quality-assurance module.
What it does:
- opens the output HDF5 file
- checks for expected datasets
- computes summary metrics such as:
- packet counts
- ADC statistics
- timestamp range
- light waveform counts if present
- creates quick plots for validation
Why it matters:
- this is the first layer of output validation
- it helps answer the question: "does this run look reasonable?"
This module handles reproducibility metadata.
What it does:
- records timestamps and environment information
- records selected environment variables
- collects git information for the
larnd-simcheckout - helps build
manifest.jsonfor each run
Why it matters:
- reproducibility is a major part of the workflow
- this file makes each run easier to understand and reproduce later
This is the profiling helper module.
What it does:
- finds profiling outputs, especially
nsysresults - runs summary commands such as
nsys stats - saves a simplified JSON report
Why it matters:
- raw profiling outputs can be hard to read directly
- this module makes them easier to compare between runs
This file handles sweep-related configuration logic.
What it does:
- reads YAML files
- loads sweep variants from
configs/sweep.yaml - writes updated YAML if needed
Why it matters:
- the sweep workflow depends on a clean way to define multiple run variants
This file contains small helper utilities used across the project.
Typical examples include:
- creating directories safely
- reading and writing JSON
- collecting timestamps
- merging dictionaries
Why it matters:
- it keeps repeated helper logic out of the main workflow files
This file defines the parameter sweep used by sim2spec sweep.
What it does:
- lists four named variants, each with a distinct random seed
- the seed drives stochastic variation so each variant produces observably different outputs
Why it matters:
- this is how multiple runs are compared in a controlled, reproducible way
This folder stores input files used for the workflow.
For this project, it is expected to contain:
MiniRun5_1E19_RHC.convert2h5.0000123.EDEPSIM.hdf5
Why it matters:
- keeping the input in a predictable place makes the notebook and shell scripts easier to follow
This folder contains two types of scripts:
- Batch scripts (
sbatch_day*.sh) — submit GPU jobs to Slurm directly withsbatch - Python helper scripts — called from the daily exercises to compare metrics, extract provenance, and save CSV outputs:
save_metrics_csv.py— Day 2: save QA metrics to CSVcompare_sweep_metrics.py— Day 3: compare packet counts and ADC stats across sweep variantsextract_sweep_provenance.py— Day 3: print seed, config, and git commit from each manifestsave_sweep_comparison_csv.py— Day 3: save full sweep comparison table tocomparison.csvcompare_profile_runs.py— Day 4: compare wall time and output file size between profile runscompare_baseline_vs_sweep.py— Day 5: compare baseline and sweep metrics side by side
Why it matters:
- keeps inline Python out of the README and Jupyter cells
- makes each step runnable with a single
python scripts/<name>.pycall
This is the lightweight environment setup script.
What it does:
- unloads any conflicting Python module
- loads
python/3.11 - defines the virtual-environment name
- prepares shell variables used by
install.sh
Why it matters:
- this is the recommended first command to source when starting a terminal session for the project
This is the main installation script for terminal users.
What it does:
- creates a virtual environment
- installs Python dependencies
- installs
sim2spec - clones and installs
larnd-sim(skips clone if already present)
Why it matters:
- this is the easiest way for terminal users to get started without following the full notebook
Standalone script for producing validation plots directly from an output HDF5 file.
python plot_validation.py "$OUTBASE/day2_baseline/run/output.h5" --outdir "$OUTBASE/day2_baseline/run/validation_plots"Why it matters:
- a quick way to visualise any run output without running the full QA pipeline
This file is the top-level overview of the repository.
Why it matters:
- it gives a quick map of the project
- it points students to the detailed guides and notebooks
These are the focused day-by-day project guides for terminal users.
What they contain:
- Day 1: environment setup, install, and smoke test
- Day 2: baseline run, QA, and validation plots
- Day 3: parameter sweeps, provenance tracking, and first profiling with Nsight Systems
- Day 4: one measurable improvement (TPB) and kernel profiling with Nsight Compute
- Day 5: final cross-day comparison and summary
Why they matter:
- each file keeps one bootcamp work block short, focused, and easier to follow during project time
The main student-facing notebook. Mirrors the Project_1_ReadMe.md through Project_5_ReadMe.md day guides with executable cells and srun-based GPU dispatch so students can run everything from inside JupyterHub.
Why it matters:
- the recommended path for students who prefer notebooks over the terminal
- keeps the notebook path shorter than the readmes while producing the same core outputs
If you are new to the repository, a good order is:
- README.md
- Project_1_ReadMe.md through Project_5_ReadMe.md
- sim2spec_perlmutter_bootcamp.ipynb — if you prefer notebooks
- src/cli.py
- src/runner.py
- src/qa.py
- src/provenance.py
- configs/sweep.yaml
Bootcamp and computing resources:
Simulation and detector workflow resources:
- 2x2 Demonstrator literature record
- DUNE larnd-sim documentation
- DUNE/larnd-sim GitHub repository
- LBL neutrino larnd-sim example
- DUNE 2x2_sim GitHub repository
- Tutorial on running 2x2_sim, April 2024
Tools used by this workflow:
- Python documentation
- Numba documentation
- CuPy documentation
- h5py documentation
- Matplotlib documentation
- Slurm workload manager documentation
- NVIDIA Nsight Systems
- NVIDIA CUDA Toolkit documentation
- This wrapper does not replace
larnd-sim; it organizes and validates runs around it. - The project is designed for Perlmutter-style Python 3.11 + venv usage.
- For Jupyter on Perlmutter, we will use the available Python environment and corresponding Jupyter kernel. Select
NERSC Pythonas the notebook kernel.