Skip to content

chore(devops): arm64 AMIs, bestool alertd and Munin on Tupaia servers - #7108

Draft
passcod wants to merge 11 commits into
devfrom
ccr-7d02eab0-4bqh4i
Draft

passcod wants to merge 11 commits into
devfrom
ccr-7d02eab0-4bqh4i

Conversation

@passcod

@passcod passcod commented Oct 6, 2026 •

Copy link
Copy Markdown
Member

Issue #:

None yet.

Changes:

  • Deployment Lambda moves out: it now lives in beyondessential/ops (pulumi/tupaia/infra/lambda, in the tupaia-infra stack), because one Lambda serves every deployment.
  • New deployment-aws/startup.sh: what a new server does once the Lambda's boot script has checked out its branch: prompt, preaggregation, Tailscale, build, pm2, nginx and observability. It's versioned with the branch. Branches without it fall back to the old steps.
  • Tailscale identity carries over on redeploy:
    • flushToS3.sh --final uploads the node's state and rejoins as an ephemeral <name>-retiring node. It never runs tailscale logout, which would invalidate the uploaded identity.
    • connectTailscale.sh restores that state only when the Lambda tagged the new instance TailscaleIdentity=handed-over, which it does only after a successful handover. The deployment then keeps the same node: same name, address and Canopy binding.
    • Otherwise, or if the restore fails, it starts clean as a new node, as before.
  • setup.sh:
    • targets Ubuntu 26.04 on our minimal BES AMIs;
    • drops libappindicator3-1 and uses libgcc-s1;
    • adds cloud-utils and the AWS CLI;
    • bakes in bestool and Munin.
  • setupObservability.sh:
    • restores Munin history;
    • installs the S3 flush (hourly, at shutdown, and on request from the Lambda);
    • serves Munin on the tailnet;
    • imports the Canopy registration and starts alertd where one exists.
  • flushToS3.sh:
    • uploads Munin history and logs;
    • after the final flush, uploads logs only, so it can't overwrite the replacement's Munin.
  • Removed: the Lambda code and setupGoldMaster.sh, both now in the ops stack.
  • Docs: the devops README and script comments point at the ops stack.

Related: beyondessential/ops#332, beyondessential/bestool#954, beyondessential/tailscale#39.

Still to do or verify:

  • Apply the ops stack. It updates the Lambda, creates the retiring key and the per-deployment roles.
  • Smoke-test PDF export on a t4g; puppeteer's Chrome on linux-arm64 is unverified.
  • Smoke-test on the 26.04 base:
    • a fresh deploy;
    • a redeploy, checking the node keeps its name and address and the old box shows up as -retiring;
    • a teardown, checking the node disappears.

Screenshots:

N/A


🦸 Review Hero

  • Run Review Hero
  • Auto-fix review suggestions
  • Auto-fix CI failures
  • Save suppressions

🤖 Generated with Claude Code

https://claude.ai/code/session_01Tzq5SfWsJpLpoBxsSG6MVg

claude added 3 commits October 6, 2026 23:49
- Deployment Lambda picks the newest gold master of the instance type's
  architecture, so x86_64 and Graviton instance types both work; cloning
  refuses to cross architectures since it copies the root volume.
- setup.sh installs bestool (bes-tools apt repo) and Munin into the AMI.
  The Image Builder pipelines now live in the `tupaia` Pulumi stack in
  ops, which replaces setupGoldMaster.sh.
- setupObservability.sh, run at boot: restores Munin history from S3 by
  deployment name and serves it on the tailnet (node:4950 and
  svc:munin-tupaia-<deployment>), and for deployments with a Canopy
  registration in Parameter Store imports it and starts alertd.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tzq5SfWsJpLpoBxsSG6MVg
…AMIs

Drops libappindicator3-1 (gone since 24.04), uses libgcc-s1, installs
cloud-utils and the AWS CLI when missing, and keeps Munin's RRDs in a
nodatacow btrfs subvolume like the Tamanu servers.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tzq5SfWsJpLpoBxsSG6MVg
@passcod passcod changed the title env(devops): arm64 AMIs, bestool alertd and Munin on Tupaia servers chore(devops): arm64 AMIs, bestool alertd and Munin on Tupaia servers Oct 6, 2026
claude added 6 commits October 7, 2026 00:08
flushToS3.sh uploads Munin history and logs (deployment, pm2, nginx,
cloud-init) to the server bucket. It runs hourly, at shutdown or
termination through tupaia-flush.service, and from the deployment Lambda,
which asks the old instance through SSM Run Command before launching its
replacement so that one restores the latest Munin data. The flush is best
effort and never blocks a redeploy.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tzq5SfWsJpLpoBxsSG6MVg
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tzq5SfWsJpLpoBxsSG6MVg
Replaces zipping and uploading by hand through the console. The function's
configuration is in the tupaia Pulumi stack in ops; this only updates its
code, including startupTupaia.sh.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tzq5SfWsJpLpoBxsSG6MVg
… flush

The Lambda's pre-replacement flush passes --final, which marks the server
superseded once the upload succeeds. Its later hourly and termination
flushes then upload logs only, so they can't overwrite the replacement's
newer Munin history.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tzq5SfWsJpLpoBxsSG6MVg
…uild in startup.sh

The Lambda (and its deploy workflow) now lives in the tupaia Pulumi stack in
beyondessential/ops. What a new server does once it has checked out its
branch moves into deployment-aws/startup.sh, versioned with the branch:
prompt, preaggregation, Tailscale, build, pm2, nginx and observability.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tzq5SfWsJpLpoBxsSG6MVg
…yment infrastructure

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tzq5SfWsJpLpoBxsSG6MVg

passcod commented Oct 7, 2026

Copy link
Copy Markdown
Member Author

The build check (CD Package Android) failed on 3b69212, but not because of this PR. Its "Fetch environment variables" step got a 503 Service Unavailable from Bitwarden Secrets Manager while authenticating. This PR changes neither the Android app nor that workflow. I've re-run the failed job once, and there is no fix to bring in.


Generated by Claude Code

claude added 2 commits October 7, 2026 00:33
…oyment>/registration

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tzq5SfWsJpLpoBxsSG6MVg
…ment

The final flush before a redeploy uploads the node's state, then rejoins the
tailnet as an ephemeral <name>-retiring node: reachable if the swap goes
wrong, gone once terminated. It never logs out, which would invalidate the
identity. The replacement takes the node over (same name, address and
Canopy binding) only when the deployment Lambda tagged it
TailscaleIdentity=handed-over, which it does only after a successful
handover; otherwise it joins as a new node, as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tzq5SfWsJpLpoBxsSG6MVg
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants