Warning
Disclaimer. This is sample code for demonstration and learning. It is not a cleared clinical product, has not been validated for diagnostic use, and must not be used to make medical decisions. It deploys real, billable AWS resources (an always-on SageMaker real-time endpoint, training and Processing jobs, AWS WAF and more) that cost money for as long as they exist. Review Costs, Cleanup and SECURITY.md before you deploy, and deploy only into a non-production account you control.
This sample is an end-to-end MLOps pipeline for binary image classification on Amazon SageMaker, written entirely in Terraform. The worked example classifies breast histopathology images (the public BreakHis dataset) as benign or malignant, but nothing in the architecture is specific to that dataset: change the data and the labels to reuse it for another binary image problem.
It covers the whole lifecycle:
- event-driven data validation and preprocessing with a patient-grouped split,
- three transfer-learning models trained in parallel and combined into an ensemble whose decision threshold is tuned on the validation split,
- a clinical quality gate and a Fairlearn fairness gate scored on the test split, with a failed gate stopping the pipeline,
- governed deployment: every model registers as
PendingManualApproval, and a person's approval triggers a blue/green endpoint update that sends 10% of the new fleet's capacity traffic first (canary) and rolls back automatically on the error and latency alarms, - a public API behind AWS WAF and an API key, with optional Amazon Bedrock reasoning for low-confidence predictions,
- scheduled drift (PSI) and fairness monitoring that can start retraining.
Everything that persists (endpoint, endpoint configuration, auto-scaling, alarms, schedules, registry) is a Terraform resource. Lambda functions only orchestrate: they serve predictions, roll an approved model onto the endpoint, and refresh the endpoint weekly.
- Architecture
- What each stack deploys
- Repository structure
- Prerequisites
- Quick start
- Deploy through CI/CD
- Call the API
- Monitoring and retraining
- Responsible AI
- Costs
- Cleanup
- Troubleshooting
- Security
- Contributing
- License
Hospital PACS is shown as the upstream image source (the sample takes JPEG or PNG uploads to S3), the validation Lambda is the upload quarantine Lambda, and Amazon VPC, Amazon Cognito and AWS CloudTrail are optional settings.
- A data batch lands in the image data bucket. When the upload finishes, a
.batch_completemarker object makes an Amazon EventBridge rule start the SageMaker pipeline. - SageMaker Pipelines validates, preprocesses, trains, evaluates and gates the model, writes every artifact under a prefix for that execution, and registers a passing ensemble in the SageMaker Model Registry as pending.
- A reviewer approves a version. The approved model is rolled onto the SageMaker real-time endpoint.
- Users open the web UI through Amazon CloudFront (static files in a private S3 bucket) and send images to Amazon API Gateway through AWS WAF with an API key. An AWS Lambda function calls the endpoint and, for low-confidence results, optionally asks Amazon Bedrock for a guarded explanation.
- The endpoint captures its responses to S3. EventBridge Scheduler runs a drift job every hour and a fairness job every day; both publish metrics to Amazon CloudWatch, whose alarms notify Amazon SNS and can start a retraining run.
- The optional CI/CD stack deploys the same stacks with AWS CodePipeline and AWS CodeBuild, and a CodeBuild project rebuilds the patched inference image in Amazon ECR every month.
- Across all stacks, customer managed AWS KMS keys encrypt data at rest, IAM roles are scoped to the project and bounded by a permissions boundary, and an AWS CloudTrail trail can be turned on for audit.
The diagrams are SVG files in docs/diagrams/; PNG exports of the same figures are in docs/diagrams/png/.
The sample is four Terraform root modules (stacks) plus reusable modules in
modules/. They are applied in this order:
| Order | Stack | Purpose |
|---|---|---|
| 1 | stack-backend-setup | Bootstrap with local state: S3 state bucket, state KMS key, workload permissions boundary |
| 2 | stack-training | Data and model buckets, SageMaker pipeline, Model Registry, Model Card, budget, optional MLflow and CloudTrail |
| 3 | stack-inference | Endpoint, API, web UI, auto-deploy, drift and fairness monitoring, Bedrock guardrail |
| 4 (optional) | stack-cicd | CodePipeline and CodeBuild that apply stacks 2 and 3 from your GitHub repository |
Creates the versioned, KMS-encrypted Terraform state bucket (S3 native state
locking, no DynamoDB), its KMS key, and the <project_name>-workload-boundary
IAM permissions boundary that every role in the other stacks carries. It keeps
its own state locally and prints the backend.hcl the other stacks use.
Creates a project KMS key, six S3 buckets (raw data, processed data, scripts,
model artifacts, inference results, monitoring) plus an SBOM bucket by default,
the SageMaker execution role, the Model Package Group, a Model Card, an AWS
Budgets budget, CloudWatch dashboards, the upload quarantine Lambda, and the
SageMaker pipeline with its EventBridge trigger. A managed MLflow tracking server (enable_mlflow) and a
multi-region CloudTrail trail (enable_cloudtrail) are available but off by
default.
- As each image lands under
medical_image_data/, the upload quarantine Lambda (enable_upload_quarantine, on by default) checks it: an allowed extension (.jpg,.jpeg,.png), opens with Pillow as that format, and at least 112 px on each side (quarantine_min_image_size_px). A file that fails is moved toquarantine/<original key>with a<original key>.reason.jsonnext to it, and the<project_name>/DataQualityQuarantinedImagesmetric counts it. The quarantine prefix is outside the prefix the pipeline reads. - Uploading the
.batch_completemarker to the data bucket fires the EventBridge rule, which starts an execution withRetrainingReason=data_upload. - Validate data (Processing job) checks image format, resolution, class balance and quality.
- Preprocess resizes images to 512 x 512 and writes train, validation and test splits grouped by patient, so images of one patient never appear in more than one split.
- Train VGG16, DenseNet121 and EfficientNetV2M in parallel training jobs with two-phase transfer learning. Training images get flips, small shifts and rotations, and a random H&E stain jitter, because stain colour differs more between patients than between classes and the models otherwise learn colour instead of tissue structure. Training jobs run with network isolation by default and read ImageNet weights from the scripts bucket.
- Evaluate tunes each model's threshold on the validation split and reports its metrics on the test split.
- Ensemble combines the three models, tunes the ensemble threshold on the validation split and scores the test split at that threshold.
- Fairness check runs Fairlearn over the ensemble's own test predictions at its tuned threshold, per magnification subgroup.
- Condition requires accuracy >= 0.85, recall >= 0.95, precision >= 0.80,
AUC >= 0.90 and a fairness disparity <= 0.10. If every check passes, the
ensemble is registered as
PendingManualApproval; otherwise the execution ends in a Fail step and nothing is registered. - Evaluation, ensemble and fairness outputs are written under
<prefix>/<pipeline-execution-id>/, so each model package points at its own weights and reports, and approving an older version deploys that version.
Creates the patched inference image (built once at apply time, then monthly), the SageMaker model, endpoint configuration and endpoint (Terraform-managed, with blue/green canary deployment, auto-rollback alarms and auto-scaling), the inference Lambda and its Pillow/NumPy layer, the API Gateway REST API with a usage plan, API key and WAF web ACL, the CloudFront web UI, the auto-deploy Lambda, the drift and fairness schedules and alarms, the retraining rule, SNS alerts, a CloudWatch dashboard, a weekly endpoint refresh, and the Bedrock guardrail when hybrid inference is on.
- The browser loads the static web UI from CloudFront, which reads it from a private S3 bucket.
- The browser posts the image directly to API Gateway, sending the API key in
the
x-api-keyheader. AWS WAF inspects the request first. - API Gateway invokes the inference Lambda, which validates and preprocesses
the image and calls the endpoint with its request id as
InferenceId. - When hybrid inference is on and the confidence is below the cutoff, the Lambda asks Amazon Bedrock for a natural-language explanation, filtered by a Bedrock guardrail.
- The endpoint writes data capture to the monitoring bucket (model output by
default; request payloads too when
data_capture_input = true). - EventBridge Scheduler starts the PSI drift job hourly and the Fairlearn job
daily. They read captured predictions, and the fairness job joins them with
the confirmed outcomes a clinician uploads to the
ground-truth/prefix. - Both jobs publish metrics to CloudWatch, and CloudWatch alarms notify the SNS alerts topic.
Optional. Creates a CodePipeline with five stages (Source, Training-Deploy, BaselineModelCreation, Script-Upload, Inference-Deploy with an optional manual approval), four CodeBuild projects, an artifacts bucket, an SNS topic for approvals and an AWS CodeConnections connection to your GitHub repository. See Deploy through CI/CD.
.
|-- stack-backend-setup/ # State bucket, state KMS key, permissions boundary (local state)
|-- stack-training/ # Buckets, upload quarantine, SageMaker pipeline, registry, Model Card, budget
|-- stack-inference/ # Endpoint, Lambda, API Gateway, WAF, CloudFront, monitoring
|-- stack-cicd/ # Optional CodePipeline and CodeBuild
|-- modules/ # Reusable Terraform modules (terraform-aws-*)
|-- scripts/ # Code the pipeline and scheduled jobs run, plus mlops_common
|-- ops-scripts/ # Local helpers: layer build, API test, dataset download, cleanup
|-- tests/ # pytest suite
|-- docs/diagrams/ # Architecture figures (SVG, PNG exports in png/)
|-- .github/workflows/ # GitHub Actions CI (static checks)
|-- Makefile # Deploy, destroy and local check targets
`-- pyproject.toml # Ruff, pytest and dev dependency group
AWS account
- An AWS account you can deploy to, with credentials that can create IAM roles, KMS keys, S3 buckets, SageMaker, Lambda, API Gateway, WAF, CloudFront and CodeBuild resources. Use a dedicated non-production account.
- Deploy in one Region (the stacks default to
us-east-1).
Service quotas
- SageMaker training: the shipped
stack-training/terraform.tfvarstrains onml.c5.2xlarge(CPU). To train on GPU (ml.g*orml.p*, which selects the GPU image automatically), request the matching "for training job usage" quota in Service Quotas first; new accounts often have 0. - SageMaker Processing:
ml.m5.4xlarge,ml.m5.largeandml.t3.mediumfor processing job usage. - SageMaker hosting: two
ml.m5.xlargefor endpoint usage (up to three with the default auto-scaling maximum), plus room for the second fleet during a blue/green update: a rollout runs the old and new fleets side by side, so allow at least five.
Amazon Bedrock
- Hybrid inference is off in the module default but on in the shipped
stack-inference/terraform.tfvars(enable_bedrock_hybrid_inference = true). When it is on, the model inbedrock_model_id(defaultus.amazon.nova-pro-v1:0, a cross-Region inference profile) must be available to the account. Set it tofalseif you do not want Bedrock.
Tools
| Tool | Version | Used for |
|---|---|---|
| Terraform | ~> 1.15 (every stack's versions.tf) |
All stacks |
| AWS CLI | v2 | local-exec build of the patched image, approvals, cleanup |
| Python | 3.13 or later (pyproject.toml) |
Helper scripts and tests |
| uv or pip | uv recommended | Building the Lambda layer, installing tools |
| make, bash, zip | any recent | Makefile targets, layer build |
make seed-baseline and scripts/download_pretrained_weights.py build Keras
models locally, so the Python that python3 resolves to needs boto3 and
tensorflow==2.19.0 (the version the endpoint serves). A virtual environment
is the simplest way to keep that separate. ops-scripts/test_api.py needs
requests and Pillow. For the local checks you also need ruff, pytest
and checkov (pip install --group dev installs ruff and pytest).
Run everything from the repository root. The Makefile honours
PROFILE=<name> (or an exported AWS_PROFILE); every deploy target runs
terraform plan -out and applies the saved plan.
0. Configure. Review stack-training/terraform.tfvars and
stack-inference/terraform.tfvars: project_name, environment,
aws_region, budget_alert_emails, alert_email, instance types and the
Bedrock toggle. training_data_path must match the prefix
scripts/data_uploader.sh writes to (medical_image_data/).
export AWS_PROFILE=<your-profile>
aws sts get-caller-identity1. Bootstrap the state backend.
make bootstrap2. Write backend.hcl for the other stacks.
make backend-config3. Build the Lambda layer. Optional: stack-inference builds the zip at
apply time when it is missing.
make layer4. Deploy the training stack.
make deploy-training5. Upload ImageNet weights (network isolation on, the default). Training
jobs have no internet access, so the trainers read weights from the scripts
bucket. Do this once before the first pipeline run. The target does nothing
when enable_network_isolation = false; it needs TensorFlow locally.
make weights6. Upload the scripts and seed the baseline model.
make seed-baselineThe endpoint cannot exist without an approved model, so this registers a
placeholder baseline (a toy network that cannot classify images) as
PendingManualApproval and prints its ARN. It does nothing when the group
already holds an approved or pending package.
7. Approve the baseline. Nothing is approved automatically; this is a deliberate manual step.
GROUP=$(terraform -chdir=stack-training output -raw model_package_group_name)
ARN=$(aws sagemaker list-model-packages --model-package-group-name "$GROUP" \
--sort-by CreationTime --sort-order Descending --max-results 1 \
--query 'ModelPackageSummaryList[0].ModelPackageArn' --output text)
aws sagemaker update-model-package --model-package-arn "$ARN" \
--model-approval-status Approved \
--approval-description "Placeholder for first endpoint deploy"8. Deploy the inference stack. The target looks up the newest Approved
package in the group, the same query the CI/CD Inference-Deploy stage runs;
set TF_VAR_model_package_arn to pin a different version.
make deploy-inferenceThe first apply also builds the patched inference image in CodeBuild and waits for it, which takes several minutes.
make deploy-all runs steps 1 to 6 and 8 in order. It cannot approve the
baseline for you, so on a first run it stops before the inference stack with
No Approved model package in <group>; approve the baseline (step 7) and run
make deploy-inference.
9. Train a real model. Download BreakHis (about 4 GB) and upload it; the
uploader writes the .batch_complete marker that starts the pipeline.
./ops-scripts/data_download_breakhis.sh data/breakhis
./scripts/data_uploader.sh data/breakhisThe upload quarantine Lambda checks every file as it lands and moves bad ones
to quarantine/ in the raw data bucket, each with a .reason.json
(stack-training README). The
pipeline's validation step checks the images again, so a file the Lambda has
not reached when the marker starts the run still fails validation.
Expect the gates to fail on BreakHis with the default models. BreakHis has 82 patients, so the patient-grouped test split holds 13 of them (4 benign), and a single patient the models get wrong moves recall or precision by several points. In a September 2026 run the ensemble scored accuracy 0.76, recall 0.90, precision 0.76 and AUC 0.79 on the test split (AUC 0.96 on validation), so the pipeline stopped at the Fail step and registered nothing. Most errors came from three test patients whose stain colour looks like the other class. Treat BreakHis as a way to exercise the pipeline; a model that clears the gates needs more patients or a stronger model, not lower thresholds.
When the execution registers a version, check its metrics in the registry and
approve it with the same update-model-package command. The approval event
triggers the auto-deploy Lambda, which rolls the new version onto the endpoint
with a blue/green canary update. Keep the file names: the patient-grouped split and
the fairness gate read the BreakHis patient id and magnification from them.
- A push to the tracked branch of your GitHub repository starts the pipeline through AWS CodeConnections.
- Training-Deploy installs a pinned Terraform release, verifies it against
HashiCorp's
SHA256SUMS, and appliesstack-trainingfrom a saved plan. - BaselineModelCreation registers the placeholder baseline as
PendingManualApprovalon the first run and prints its ARN in the build log. - Script-Upload uploads the pipeline scripts and the ImageNet weights to the scripts bucket.
- Inference-Deploy waits for the manual approval action
(
require_manual_approval, defaulttrue), then builds the Lambda layer and appliesstack-inferencewith the newest Approved package in the group. It fails if no package is approved yet. - Both Terraform stages keep their state in the KMS-encrypted state bucket
created by
stack-backend-setup.
Steps:
- Push this repository to a GitHub repository you own.
- Run
make bootstrapandmake backend-config(the pipeline uses the same state bucket). - In
stack-cicd/terraform.tfvars, setgithub_owner,github_repo,github_branchandstate_bucket_name(printed bymake backend-config). They have no usable defaults and the placeholders are rejected. - Run
make deploy-cicd. - Complete the CodeConnections handshake once: in the console open
Developer Tools > Settings > Connections, select
<project_name>-github, choose Update pending connection, and install the AWS Connector for GitHub app on the account that owns your repository. Check thatterraform -chdir=stack-cicd output github_connection_statusshowsAVAILABLE. - Run the pipeline: choose Release change in CodePipeline or push to the tracked branch.
- On the first run, approve the baseline package (ARN in the BaselineModelCreation log, command as in Quick start step 7) before you approve the pipeline's manual approval action.
Subscribers in approval_notification_emails must confirm the SNS
subscription before they get approval notices.
The API requires an API key in the x-api-key header. Terraform outputs the
URL and the key:
API_URL=$(terraform -chdir=stack-inference output -raw api_gateway_url)
API_KEY=$(terraform -chdir=stack-inference output -raw api_key_value)
python3 ops-scripts/test_api.py --api-url "$API_URL" --api-key "$API_KEY" \
--image-path data/breakhis/breast_benign/<file>.pngWithout --image-path the script samples --count images per class from
--data-path (default data/) and checks each prediction against the folder
label. Without --api-url and --api-key it reads both from terraform output.
The API has two routes:
POST /predictwith a JSON body{"image": "<base64 image>"}. The response carriesrequest_id,prediction,confidence,probabilities,threshold_used,routingand, when hybrid inference answers,reasoning. Images larger thanmax_image_bytes(5 MB by default) return HTTP 413.GET /results/{request_id}returns the stored prediction for that request.
Add "explain": true to the POST /predict body to get per-image
explanations with the prediction (--explain in test_api.py). They are off
unless asked for, so a request without the field gets the same response, at
the same latency, as before. {"explain": {"methods": ["gradcam"]}} asks for
one method only (gradcam or region_shapley).
{
"request_id": "3f1c2a9e-5b7d-4e2a-9c1f-0a1b2c3d4e5f",
"prediction": "malignant",
"confidence": 0.91,
"probabilities": {"benign": 0.09, "malignant": 0.91},
"threshold_used": 0.42,
"model_info": "Image analysis using ensemble model",
"routing": "high-confidence-custom",
"explanations": {
"gradcam": {
"method": "grad-cam",
"target": "malignant",
"layer": "feature map feeding global average pooling (last convolutional block)",
"combination": "ensemble-weighted average of per-model maps, each scaled to max 1",
"grid_size": 14,
"grid": [[0.021, 0.054, 0.112]],
"top_region": {"box": [0.2857, 0.1429, 0.6429, 0.5], "peak_cell": [4, 6], "cells": 19},
"elapsed_ms": 310
},
"region_shapley": {
"method": "region Shapley values estimated by antithetic permutation sampling",
"target": "malignant probability",
"grid_size": 4,
"values": [[0.012, -0.003, 0.041, 0.0]],
"baseline": "ImageNet mean colour (0 after normalisation)",
"baseline_score": 0.3712,
"full_score": 0.91,
"sum": 0.5388,
"permutations": 2,
"evaluations": 32,
"seed": 0,
"elapsed_ms": 4100
},
"explain_ms": 4420
}
}grid (14 x 14) and values (4 x 4) are cut to their first row here. The
whole explanations object is under 3 KB. How they are computed, on the
endpoint, for the ensemble that made the call:
- Grad-CAM for the predicted class. Each member model's SavedModel carries
a
gradcamsignature: activations are the feature map that feeds its global average pooling layer (the last convolutional block), gradients are of the malignant logit (negated for a benign call), and the map is resized to 14 x 14. The handler scales each member's map to a maximum of 1 and averages them with the ensemble weights.top_region.boxis[x_min, y_min, x_max, y_max]as fractions of the image, around the connected cells at 50% of the peak or more. - Region Shapley values, estimated by sampling. The image is split into a
4 x 4 grid; a region left out is set to the ImageNet mean colour. The value
of a region is its average marginal effect on the ensemble's malignant
probability over random orderings of the regions, drawn in antithetic
pairs with a fixed seed. The values always sum to
full_score - baseline_score. This is not theshaplibrary, and 32 evaluations of 16 regions is a coarse estimate, not an exact Shapley value.
Each explained request costs up to explanation_max_evaluations (default 32)
masked-image passes through every member model, within
explanation_time_budget_ms (default 5000). When time runs short, fewer
permutations run and permutations says how many. Set
enable_request_explanations = false to turn the option off. A model package
trained before this feature reports "status": "unavailable"; retrain to get
the signatures.
The web UI is at terraform -chdir=stack-inference output -raw frontend_cloudfront_url. It carries the same API key, so the key identifies
and meters callers; it does not authenticate them (see SECURITY.md).
The figure leaves out the review step: a reviewer approves the new version in the Model Registry before the auto-deploy Lambda rolls it out, and the fairness alarm starts retraining the same way as the drift alarm.
- A drift or fairness alarm in CloudWatch changes state to ALARM.
- An EventBridge rule starts the SageMaker pipeline with
RetrainingReason=drift_detected. - A version that clears the gates is registered as
PendingManualApprovaland waits for a reviewer. - The reviewer approves the package in the Model Registry.
- The approval state change triggers an EventBridge rule that invokes the auto-deploy Lambda.
- The Lambda creates a model and endpoint configuration for the approved
package and updates the endpoint with the stack's deployment policy.
SageMaker starts a new (green) fleet and, by default, routes 10% of its
capacity (
canary_size_percent, rounded up to one of the two instances) to it fortraffic_shift_wait_intervalseconds (300). If neither the error alarm (5XX or model errors) nor the latency alarm fires, the rest of the traffic shifts, the old fleet is kept fortermination_wait_seconds(120) more and then terminated. An alarm at any point rolls all traffic back to the old fleet. The weekly endpoint refresh uses the same policy.traffic_routing_typeselectsCANARY(default),LINEAR(linear_step_percentper step) orALL_AT_ONCE; canary and linear need at least two instances (endpoint_initial_instance_countandendpoint_min_capacity, both 2 by default), and serverless endpoints always shift all at once. - Failed asynchronous invocations of the Lambda go to an Amazon SQS dead-letter queue, and a queue-depth alarm notifies the SNS alerts topic.
What the monitors measure:
- Drift (
scripts/drift/compute_drift.py, hourly): the Population Stability Index of the endpoint's malignant-probability scores over the last 24 hours against a baseline score distribution. The alarm threshold isdrift_threshold(0.2 by default, 0.1 in the shipped tfvars). It publishes nothing with fewer than 30 captured predictions or without a baseline. The auto-deploy Lambda writes the baseline after each rollout: it copies the deployed package's test scores (predictions.jsonnext to itsmodel.tar.gz) tos3://<monitoring-bucket>/monitoring/baselines/output-only/statistics.json. The placeholder baseline has no test scores, so drift is reported only once a trained model has been approved and rolled out. - Fairness (
scripts/fairness/compute_fairness.py, daily): joins captured predictions with confirmed outcomes uploaded as JSON Lines ({"request_id": ..., "label": 0|1, "group": "<subgroup>"}) to theground-truth/prefix of the monitoring bucket, and publishes the larger of the demographic-parity and equalized-odds differences. The sample does not populate that prefix; until someone does, the job publishes no metric. - Operations: API Gateway 4XX/5XX, endpoint errors and latency alarms, a
composite health alarm, and a CloudWatch dashboard. Set
alert_emailto get the SNS notifications.
ops-scripts/ops_generate_endpoint_traffic.py sends labelled images straight
to the endpoint to seed data capture before real traffic exists, and
ops-scripts/verify_monitoring.sh checks that the alarms, topics and rules
exist. The drift and fairness jobs need data capture, so they are skipped when
use_serverless_inference = true.
In the pipeline one condition step applies both gates after the fairness check, and a failing execution ends in a Fail step with its evaluation and fairness reports kept under its execution prefix.
- Fairness gate in the pipeline. Fairlearn scores the ensemble that would be registered, at its tuned threshold, per subgroup. BreakHis carries no demographic data, so the subgroup is image magnification, an honest proxy. With fewer than two subgroups the result is "not evaluable" and the gate fails. Supply a real sensitive attribute before drawing any fairness conclusion.
- Fairness on live traffic, held to the same bar, once confirmed outcomes are uploaded (see above).
- Human approval before any model serves traffic, including the baseline and drift-triggered retrains.
- Model Card recording intended use, a High risk rating and clinical
caveats (
enable_model_card). - Per-request explanations (opt-in per call):
"explain": truereturns a Grad-CAM heatmap and region Shapley values with the prediction, computed on the endpoint for the served ensemble (see Explanations in the response). They show which image regions drove the score. They are decision support for a reviewer and are not validated against pathologist annotations. - Hybrid inference (optional): low-confidence predictions get a Bedrock explanation constrained by a guardrail (content filters, name anonymization).
- Human review with Amazon A2I (optional,
enable_human_review, off by default). Amazon A2I is in maintenance mode and not open to new customers, so this only works in accounts that already use it.
You pay for the resources while they exist. There are no NAT gateways, AWS
Network Firewall or VPC endpoints by default (nothing runs in a VPC unless you
set vpc_config in stack-training). Estimate your configuration with the
AWS Pricing Calculator.
| Resource | Billed | Notes |
|---|---|---|
SageMaker real-time endpoint (ml.m5.xlarge, 2 to 3 instances) |
Every hour it runs | The main standing cost, about $170 per instance per month in us-east-1, so about $340 for the default two (check your Region). Two is the minimum for canary traffic shifting; with traffic_routing_type = "ALL_AT_ONCE" you can run one. Both fleets are billed while a blue/green update runs. use_serverless_inference = true scales to zero but disables data capture and monitoring |
| SageMaker training jobs | Per run | 3 jobs per execution; CPU ml.c5.2xlarge in the shipped tfvars, GPU if you change it. Managed Spot is available (enable_managed_spot_training) |
| SageMaker Processing jobs | Per run | 5 per pipeline execution, plus the hourly drift job and daily fairness job (ml.t3.medium, capped at 900 s) |
| AWS WAF | Monthly per web ACL and rule, plus per request | One regional web ACL with three rules (api_enable_waf) |
| API Gateway, Lambda, CloudFront | Per request | Low at demo traffic |
| Amazon Bedrock | Per token | Only for low-confidence predictions when hybrid inference is on |
| SageMaker MLflow tracking server | Every hour it runs | Off by default (enable_mlflow) |
| AWS CloudTrail trail | Per event delivered, plus S3 | Off by default (enable_cloudtrail); leave it off when an organization trail covers the account |
| Amazon S3 | Storage and requests | Buckets are versioned; images, capture and model artifacts accumulate |
| CodeBuild, ECR | Build minutes, storage | Monthly patched-image rebuild; the last 10 images are kept; CI/CD builds if you use stack-cicd |
| KMS, CloudWatch, SNS, SQS, EventBridge | Keys, metrics, alarms, logs | Small |
stack-training creates an AWS Budgets budget (monthly_budget_usd, default
200) scoped by the Project tag; set budget_alert_emails to be notified.
Destroy the stacks when you are not using them.
force_destroy is false on the buckets, so they must be emptied before
Terraform can delete them. make destroy does that for you:
make destroy PROFILE=<your-profile>It runs ops-scripts/cleanup_all.sh --destroy, which asks you to type the
project name, then:
- deletes what pipeline runs and Lambda functions created outside Terraform (endpoint, endpoint configurations, models, model packages, job log groups),
- empties the project's buckets, including all object versions, and never touches the Terraform state bucket,
- runs
terraform destroyinstack-cicd, thenstack-inference, thenstack-training(stack-inferencereadsstack-trainingstate, so this order matters).
Finally, empty the state bucket (it keeps state versions) and destroy the bootstrap stack:
cd stack-backend-setup && terraform destroyRun ./ops-scripts/cleanup_all.sh --help for --project, --region,
--state-bucket and --yes. Afterwards, check in the console that the
SageMaker endpoint and, if you enabled it, the MLflow tracking server are gone.
Custom CloudWatch metrics cannot be deleted and expire on their own.
| Symptom | Cause and fix |
|---|---|
Pipeline Source stage fails, or github_connection_status is PENDING |
The CodeConnections handshake is not complete. Finish it in Developer Tools > Settings > Connections (see Deploy through CI/CD), then release the change again. |
Training job fails with ResourceLimitExceeded |
No quota for the training instance type. Request the "for training job usage" quota in Service Quotas, or train on CPU (ml.c5.2xlarge, the shipped value). Instance types starting ml.g or ml.p use the GPU image; others use the CPU image. |
| Training job fails at start or cannot load weights | Network isolation is on and pretrained-weights/ in the scripts bucket is empty. Run make weights (Quick start step 5), or set enable_network_isolation = false. |
| Pipeline ends at the Fail step | A gate failed. Read clinical_quality_gate in ensemble_results.json and bias_metrics.json under the execution's prefix in the model artifacts bucket. A not_evaluable fairness status means the file names carried fewer than two magnification subgroups; for a small demo dataset set allow_not_evaluable = true in stack-training/terraform.tfvars. |
make deploy-inference stops with No Approved model package in <group>, or the plan fails with model_package_arn must be the versioned ARN of an Approved package |
Nothing in the group is approved yet, or TF_VAR_model_package_arn is not a versioned ARN. Approve the baseline (Quick start step 7) and run make deploy-inference again. In CI/CD the Inference-Deploy stage fails with no Approved model package for the same reason. |
Lambda layer zip missing: run 'make layer' |
make deploy-inference checks for the zip first. Run make layer. A plain terraform apply builds it automatically when missing (it needs bash, zip and uv or pip); delete the zip to force a rebuild. |
BucketNotEmpty on destroy |
Buckets have force_destroy = false. Use make destroy, or run ./ops-scripts/cleanup_all.sh and retry the destroy. |
| No fairness metric in CloudWatch | Expected until confirmed outcomes are uploaded to ground-truth/ in the monitoring bucket and at least 50 predictions match them. The job's CloudWatch log says why it published nothing. |
| No drift metric in CloudWatch | Fewer than 30 captured predictions in the window, or no baseline at monitoring/baselines/output-only/statistics.json yet: the auto-deploy Lambda writes it when it rolls out a trained package (see Monitoring and retraining). |
make deploy-training fails with ResourceAlreadyExistsException for /aws/sagemaker/ProcessingJobs or /aws/sagemaker/TrainingJobs |
SageMaker or an earlier deployment already created these account-wide log groups. Import them and apply again: terraform -chdir=stack-training import 'module.cloudwatch_monitoring.aws_cloudwatch_log_group.processing_jobs[0]' /aws/sagemaker/ProcessingJobs (and the same for training_jobs[0] with /aws/sagemaker/TrainingJobs), with the same TF_VAR_* values make deploy-training sets. The stack then owns them, so a destroy deletes them. |
make seed-baseline fails on import tensorflow |
Install boto3 and tensorflow==2.19.0 into the Python that python3 resolves to. |
| API returns 403 | Missing or wrong x-api-key, or the WAF rate rule blocked your IP. |
Uploaded images are missing from medical_image_data/ |
The upload quarantine Lambda moved them. Each one is under quarantine/<original key> in the raw data bucket with a .reason.json naming the failed check (extension, does not open, wrong format for its extension, or smaller than quarantine_min_image_size_px). Fix the file and upload it again. |
make deploy-inference plan fails with traffic shifting sends ... of the new fleet first |
Canary and linear traffic shifting need at least two instances. Keep endpoint_initial_instance_count and endpoint_min_capacity at 2 or more, or set traffic_routing_type = "ALL_AT_ONCE" to run one instance. |
See SECURITY.md for the security model, known limitations, what you must still do before any real use, and how to report a vulnerability.
See CONTRIBUTING.md. Before you open a pull request, run the same checks as CI (.github/workflows/ci.yml):
make fmt validate lint testThis library is licensed under the MIT-0 License. See the LICENSE file. Dataset licensing is your responsibility; confirm the terms of any dataset you use and never commit PHI or PII.