Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
20 changes: 19 additions & 1 deletion docs/upgrade.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,6 +27,12 @@ a version suffix, and the installed ate-api-server serves
install instead, because its DaemonSet selector cannot be changed in
place.

Nothing drains worker nodes on its own: node auto-upgrade is off on
every pool that runs workers and none of them is spot or preemptible,
as the
[Create Cluster warning](../tools/setup-gcp/README.md#2-create-cluster)
requires.

Actor snapshots are readable by both the old and the new build. An
actor can therefore suspend on one version and resume on the other in
either direction, which is what lets the two versions serve side by
Expand All @@ -48,7 +54,19 @@ Three things break an upgrade.
same node, and old workers end up next to the new atelet: exactly
the version skew the roll exists to prevent.
2. **Do not edit a serving worker pool.** The controller would roll
the pool's Deployment straight through live actors.
the pool's Deployment straight through live actors. A deleted
worker pod does go through the eviction path: `SIGTERM` is
forwarded into the actor's containers and the control plane keeps
accepting a suspend for about 60 seconds, so an actor suspended
inside that window saves its state and stays resumable. Handling
`SIGTERM` by exiting cleanly is not enough on its own; the suspend
has to reach the control plane and finish. An actor still awake
when the window closes moves to `ACTOR_STATE_CRASHED`, which is
terminal: `resume` and `suspend` are both refused, there is no
recover verb, and the snapshot the actor still holds cannot be
used to start it. It has to be deleted and recreated, losing its
state. The same applies to scaling a serving pool down, which
removes pods without suspending the actors on them.
3. **(If on GKE) Do not touch the node pool's label until every node
is rolled.** A pool label update applies in place to every node in
the pool, so the whole fleet flips at once, with no drain and no
Expand Down
28 changes: 28 additions & 0 deletions tools/setup-gcp/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -95,6 +95,34 @@ Filestore CSI driver disabled).
> podcertificate ClusterTrustBundles to be ready" and `kubectl get
> clustertrustbundles` reports the resource type is not served.

> [!WARNING]
Comment thread
mayawang marked this conversation as resolved.
> **Turn node auto-upgrade off on any node pool that runs workers, and do not
> use spot or preemptible nodes for them.** When a worker pod is deleted,
> `SIGTERM` is forwarded into the actor's containers and the control plane keeps
> accepting a suspend for about 60 seconds. An actor suspended inside that
> window keeps its state. One still awake when the window closes is moved to
> `ACTOR_STATE_CRASHED` with its worker assignment cleared, and `CRASHED` is
> terminal: `resume` and `suspend` are both refused, there is no recover verb,
> and the snapshot the actor still holds cannot be used to start it. The only
> way out is to delete the actor and create a new one, which loses its state.
>
> Auto-upgrade is the trigger to plan for, because GKE enables it by default and
> it fires on Google's maintenance schedule rather than yours. `create cluster`
> does not disable it, so do it yourself on every pool that runs workers:
>
> ```bash
> gcloud container node-pools update "${NODE_POOL}" \
> --cluster "${CLUSTER_NAME}" --location "${CLUSTER_LOCATION}" \
> --no-enable-autoupgrade
> ```
>
> This is a management setting, so it takes effect without recreating nodes and
> is safe to apply to a serving cluster. Node auto-repair, preemption and OOM
> kills reach the same path and cannot be configured away, so treat the setting
> as removing the scheduled risk rather than all of it. Change versions through
> the [rolling upgrade runbook](../../docs/upgrade.md), which has you suspend
> every actor on a node at your own pace before the node moves.

```bash
go run ./tools/setup-gcp create cluster [flags]
```
Expand Down
Loading