This document describes how to stage, validate, and publish a stable TorchTitan release.
The examples use TorchTitan 0.3.0, PyTorch 2.14.0, torchvision 0.29.0,
and torchao 0.18.0. Update all version numbers and wheel index URLs for each
release train.
- Follow each PyTorch minor release, approximately every two months: cut the
corresponding TorchTitan release branch shortly after PyTorch creates
release/X.Y, then publish TorchTitan with the PyTorch stable release. - For an urgent patch, publish
0.Y.(Z+1)from the existingrelease/0.Ybranch. A new PyTorch minor release is not required.
Complete both CI validation and release-specific testing before cutting the branch.
Check the latest main or scheduled run for every applicable workflow,
including:
- lint and CPU/GPU unit tests;
- 8-GPU Real-PG feature and model integration tests;
- H100 integration tests; and
- integration tests for projects under
torchtitan/experiments/, including GraphTrainer, the Transformers modeling backend, RL, and TorchFT.
Path-filtered experimental workflows might not run on every main commit, so
check their latest scheduled run explicitly. Investigate all failures before
cutting the release branch. CI must be green, but CI status alone does not
qualify a release.
cd ~/torchtitan
uv venv --python 3.12 .venv
source .venv/bin/activate
uv pip install -e . -r requirements.txt -r requirements-dev.txtInstall the PyTorch release-staging packages. Update the versions and CUDA wheel index for the release being tested.
uv pip install --force-reinstall --pre \
--index-url https://download.pytorch.org/whl/test/cu130 \
"torch==2.14.0" \
"torchvision==0.29.0" \
"torchao==0.18.0"pre-commit run --all-files
pytest tests/ -xSmoke-test the example and tutorial commands documented in the repository. Confirm that the instructions are accurate and that each selected job completes.
Use 0.Y.0rc1 for the first release candidate and 0.Y.0 for the final
release. Check the latest published version in the
PyPI release history.
git checkout main
git pull origin main
git checkout -b release/0.3
git push -u origin release/0.3Edit assets/version.txt to the RC version, then open
the PR against the release branch, not main.
Use Semantic Versioning for the base 0.Y.Z version:
- increment
Yfor a feature release, for example0.3.0->0.4.0; - increment
Zfor a patch release containing fixes, for example0.3.0->0.3.1; and - append the Python RC suffix
rcNfor release candidates, for example0.3.0rc1, then0.3.0rc2if another candidate is required.
git checkout release/0.3
git pull origin release/0.3
echo "0.3.0rc1" > assets/version.txtOpen a PR targeting release/0.3, wait for CI, and merge it.
Before staging a new release:
- In
.github/workflows/validate_rc.yaml, update the pinnedtorch,torchvision,torchao, andtritonversions. Set Triton to the version required by the selected PyTorch wheel. - In
.github/workflows/validate_release_gpu.yml, update the pinnedtorch,torchvision, andtorchaoversions and the PyTorch wheel index. Triton is installed transitively by the GPU PyTorch wheel and is verified at runtime rather than pinned separately. - Merge both workflow updates into the release branch before staging the RC.
The
test_release.yml workflow builds the
wheel and source distribution, runs twine check --strict, and uploads the
artifacts to TestPyPI. Before building, it invokes the reusable lint workflow
against all files on the selected release branch.
To run it:
- Open GitHub Actions -> Publish a Release to TestPyPI.
- Click Run workflow.
- Select the release branch, for example
release/0.3. - Start the workflow.
- Confirm that the lint and build jobs pass.
- When prompted, open Review deployments, select the protected
test-releaseenvironment, and click Approve and deploy.
Confirm that the RC appears in the TestPyPI release history. Do not proceed until the package is visible and installable.
The TestPyPI staging workflow automatically invokes the validate-rc job after
publishing. It does not run on push or on a schedule.
For Python 3.11 and 3.12, the job creates a clean virtual environment and:
- installs the pinned
torch,torchvision,torchao, andtritonpackages from the PyTorch test CPU channel; - installs the exact TorchTitan RC from TestPyPI;
- verifies the TorchTitan version and confirms that it was imported from
site-packages; - runs a short CPU forward, backward, and optimizer smoke test;
- verifies that the loss is finite and decreases; and
- runs all enabled CPU unit tests against the installed RC from an isolated temporary directory.
Confirm that every validate-rc matrix job is green.
Optional manual GPU installation check:
python -m pip install --pre \
--index-url https://download.pytorch.org/whl/test/cu130 \
"torch==2.14.0" \
"torchvision==0.29.0" \
"torchao==0.18.0"
python -m pip install \
--index-url https://test.pypi.org/simple/ \
--extra-index-url https://pypi.org/simple/ \
"torchtitan==0.3.0rc1"
python - <<'PY'
import torch
import torchtitan
assert torch.cuda.is_available()
print(f"torchtitan={torchtitan.__version__} ({torchtitan.__file__})")
print(f"torch={torch.__version__}, cuda={torch.version.cuda}")
print(f"gpu={torch.cuda.get_device_name(0)}")
PYAfter the RC is available on TestPyPI, run the Validate a TestPyPI Release Candidate on GPUs workflow from the release branch. Every job installs the exact TorchTitan RC from TestPyPI together with the pinned PyTorch release-staging packages.
To run it:
- Open GitHub Actions -> Validate a TestPyPI Release Candidate on GPUs.
- Click Run workflow.
- Select the release branch, for example
release/0.3. - Start the workflow and wait for every validation job to finish.
The workflow runs six GPU jobs: four on 8x A10G runners and two on 8x H100 runners. All integration suites use real process groups:
validate-standard-gpuwith thecoresuite checks the installed RC and dependency versions, all enabled single-GPU and multi-GPU unit tests, the Real-PG feature and model suites, their integrated loss and gradient-norm goldens, and Flux;validate-standard-gpuwith thegraph-trainersuite runs the standard GraphTrainer integrations, numerics, graph passes, profiler, tracing, precompile, bitwise-determinism, and SAC peak-memory tests;validate-standard-gpuwith thetorchftsuite installs the pinned stabletorchft==0.2.0, starts Lighthouse, and runs the 8-GPU TorchFT integration test for 10 training steps with checkpointing enabled;validate-standard-gpuwith thetransformers-modeling-backendsuite installstransformers==5.9.0and runs the MoE FSDP+TP+EP+CP, dense FSDP+TP+PP, dense CP+PP, and SFT integration tests;validate-h100runs the base H100 suite, Qwen3 with DeepEP v2, and DeepSeek V3 with HybridEP as separate tests; andvalidate-graph-trainer-h100runs the H100 GraphTrainer integrations, MoE numerics, DeepSeek V3 precompile, and bitwise-determinism tests.
Every GPU test job installs and imports the exact TestPyPI RC from
site-packages and checks the expected torch, torchvision, and torchao
versions. RL and GraphTrainer AutoParallel are intentionally not included.
Before finalizing the version, confirm that:
- the Step 4A Publish a Release to TestPyPI workflow is green, including its
Python 3.11 and 3.12
validate-rcjobs; - the Step 4B Validate a TestPyPI Release Candidate on GPUs workflow is
green, including
Validate core on 8x A10G,Validate graph-trainer on 8x A10G,Validate torchft on 8x A10G,Validate transformers-modeling-backend on 8x A10G,Validate core H100 tests, andValidate GraphTrainer H100 tests; - searching the GPU job logs for
SKIPPEDandSkipping testfinds only expected exclusions.
If the staged package needs a code fix, follow Step 5, increment the RC version, publish the new RC, and repeat the validation. Never reuse an RC version.
Identify the root cause before changing the release. The fix path depends on whether the problem is in TorchTitan, PyTorch, torchao, or the validation environment.
Fix the issue on main first, then cherry-pick the merged fix to the release
branch.
-
Land the fix PR on
mainand record its merge commit SHA. -
Cherry-pick the commit onto the release branch:
git checkout release/0.3 git pull origin release/0.3 git cherry-pick <sha> # Resolve conflicts if necessary, then push the branch. git push origin release/0.3
-
Update
assets/version.txtto the next RC, for example0.3.0rc2. -
Repeat Steps 3-4 to publish and validate the new RC. Never reuse an RC version.
If an urgent fix cannot wait for main CI, open the fix PR directly against
release/0.3, then forward-port the identical change to main.
- Land the fix in
pytorch/pytorchmain, or work with a PyTorch developer to land it. - Mark it as a release blocker for
X.Yand request a cherry-pick intopytorch/pytorchrelease/X.Y. The release manager must approve it before the cherry-pick deadline. - When PyTorch publishes the next RC, update the pinned version and repeat Steps 3-4.
- If the upstream fix cannot land in time, use a documented TorchTitan workaround and remove it after the upstream fix is available.
The published wheel cannot be changed. File the upstream issue, wait for a
PyTorch patch release such as X.Y.1, then update the pin and revalidate. If
the release cannot wait, use a documented TorchTitan workaround.
The same policy applies to torchao issues.
After all CPU and GPU RC validations are green:
-
Update
assets/version.txton the release branch:git checkout release/0.3 git pull origin release/0.3 echo "0.3.0" > assets/version.txt
-
Open a PR targeting
release/0.3. -
Confirm CI is green and merge the PR.
- Open the new release page.
- Set the tag to
v0.3.0and the target branch torelease/0.3. - Click Generate release notes. Verify that the Full Changelog compares against the previous release tag, then organize the changes and add the pinned torch and torchao versions plus a short highlight summary.
- Do not select Set as a pre-release for the final stable release.
- Click Publish. This triggers
.github/workflows/release.yml. - When prompted, approve the protected
releaseenvironment deployment.
Use a clean environment and verify the actual installed wheels:
python -m venv /tmp/torchtitan-release-verify
source /tmp/torchtitan-release-verify/bin/activate
python -m pip install --upgrade pip
python -m pip install \
--index-url https://download.pytorch.org/whl/cu130 \
"torch==2.14.0" \
"torchvision==0.29.0" \
"torchao==0.18.0"
python -m pip install "torchtitan==0.3.0"
# Run outside the repository so the checkout cannot shadow the installed wheel.
cd /tmp
python - <<'PY'
import importlib.metadata
import sysconfig
from pathlib import Path
import torchtitan
expected_version = "0.3.0"
package_path = Path(torchtitan.__file__).resolve()
site_packages = Path(sysconfig.get_paths()["purelib"]).resolve()
assert importlib.metadata.version("torchtitan") == expected_version
assert torchtitan.__version__ == expected_version
package_path.relative_to(site_packages)
print(f"torchtitan={torchtitan.__version__}")
print(f"package_path={package_path}")
PYThen:
- Confirm the release appears in the PyPI release history.
- Confirm the verification command exits successfully and prints
torchtitan=0.3.0withpackage_pathunder/tmp/torchtitan-release-verify/lib/python*/site-packages/. - Run a short debug-model training check against the installed PyPI wheel. Verify that the loss is finite and decreases. An import check alone is not sufficient.
- Pre-branch CI on
main: lint, CPU/GPU unit tests, Real-PG integration tests, H100 tests, and the latest available CI for projects undertorchtitan/experiments/. - Release-specific source validation: local unit and smoke tests plus the Real-PG feature/model suite and its integrated numerical checks against the pinned PyTorch release-staging packages.
- TestPyPI staging and CPU validation: full-repository lint runs before publication; clean Python 3.11 and 3.12 environments then install the staged RC, verify package versions and import location, run a short training step, and run the complete enabled CPU unit-test suite.
- TestPyPI GPU validation: six GPU jobs, four on 8x A10G runners and two on 8x H100 runners, install the staged RC and pinned GPU packages. All integration suites use real process groups and cover GPU unit tests, feature/model integration tests, loss and gradient-norm goldens, Flux, GraphTrainer, TorchFT, the Transformers modeling backend, and the dedicated H100 suites.
- Release scaling validation: run a representative full-scale Llama 3 405B FSDP workload with the staged RC and pinned PyTorch packages. Confirm the run completes, loss remains finite and converges, and throughput and memory show no unexpected regression.
- Production PyPI validation: a clean environment installs the final wheels, verifies the installed package, and runs a short debug-model training check.
- GraphTrainer AutoParallel: standard and H100 suites and numerics. We don’t include them because we don’t want to depend on the auto parallel main branch.
- Experimental projects: RL.
- Additional platforms: ROCm.