Skip to content

feat(hub,cli,web): fleet runner version governance (skew, self-upgrade, soft-fail reopen) - #1108

Merged
tiann merged 6 commits into
mainfrom
fix/hub-runner-version-governance
Aug 11, 2026
Merged

tiann merged 6 commits into
mainfrom
fix/hub-runner-version-governance

Conversation

@heavygee

@heavygee heavygee commented Jul 20, 2026 •

Copy link
Copy Markdown
Collaborator

Same-repo mirror of #1086 (head on tiann:fix/hub-runner-version-governance instead of the fork). Identical commits/files — opened so HAPI Bot can review while fork PR checkouts under pull_request_target are blocked pending #1107. No local tooling / operator paths in the diff.

Summary

After #1037, Cursor reopen hard-failed when the hub was newer than a remote runner: missing cursor-chat-store-status was treated as deleted chat data even when store.db still existed.

This PR makes hub↔runner generation governance explicit:

  • Soft-fail Cursor reopen when the chat-store probe errors (missing handler / skew / unknown). Definitive onDisk: false still blocks.
  • Runners advertise version + capability set on machine metadata; hub declares required capabilities (starting with Cursor chat-store probe).
  • Unmissable persistent web banner on session/machines UI listing upgrade/restart steps for skewed hosts.
  • Ensure v1: if a newer CLI binary is already on disk (installedCliMtimeMs ≠ startedCliMtimeMs), hub may call stop-runner so systemd/handoff loads the new generation; otherwise the banner stays.
  • Multi-machine install note pointing operators at the banner.
  • Machine selector shows CLI version + UPDATE REQUIRED when skewed.

Test plan

  • bun typecheck
  • bun run test (shared capability registry, soft-fail reopen, runner ensure, web reopen gate, skew banner)
  • Peer hub: live banner copy verified (Runner out of date on proxmox + upgrade/restart instructions)
  • Operator dogfood: upgrade hub first, leave remote runner old, confirm reopen is allowed with honest hint + banner; upgrade remote CLI and confirm banner clears

Issues

Fixes #1084

Upgrade notes

Fleet self-upgrade is not zero-touch on first generation or right after a hub restart.

Chicken / egg

  1. Upgrade the hub host CLI and restart the hub.
  2. On each runner host: install the matching CLI once or restart the runner so it re-registers machine RPCs. A host can look online (machine-alive) while the hub RpcRegistry is empty — every machine RPC then fails with RPC handler not registered.
  3. Wait ~30s (keepalive re-registers on current CLI generations).
  4. Probe before trusting auto-upgrade:
    • spawn JSON body must be type: "success" (HTTP 200 alone is not enough)
    • Upgrade either progresses or returns clean upgrade_unavailable — not toasting upgrade_failed when caps are advertised but the live RPC is missing
  5. True legacy runners (capabilities empty / ancient CLI) still need a manual install; the hub cannot invent runner-self-upgrade.

Guide: docs/guide/deployment.md → Fleet upgrade after a hub update.

This tip (fee9084b2)

  • Hub requires live runner-self-upgrade registration before fleet upgrade (advertised-but-not-live → upgrade_unavailable).
  • CLI re-emits rpc-register on every keepalive so a stable hub heals ghost registries without a full reconnect.
  • Hub optionally acks rpc-register.

Kill-criteria before claiming smooth fleet deploy

After hub is on this generation and each remote runner has restarted once onto it (or keepalive heal is live):

  1. Hub restart → wait 30s → without restarting the runner: spawn JSON success within one keepalive (~20s), or clean upgrade_unavailable while healing — never upgrade_failed from advertised-but-not-live caps.
  2. POST .../upgrade-runner either starts apply or returns upgrade_unavailable with restart guidance.
  3. Artifact channel: /cli/upgrade/cli-artifact must return a real binary (not SPA HTML).

Cross-session upgrade / restart nudges

Any "restart the runner", "run hapi upgrade", or remat instruction sent from one HAPI session to another must use peer delivery (ping_peer / sentFrom: peer) so the target UI shows an @session chip (verified when a session capability is present, or @name/@id with ⚠ when unattributed). Do not inject those instructions as ordinary user-composer text on the target — operators will treat that as their own keystrokes. Unverified (⚠) chips are claims only: do not auto-execute shell/upgrade actions from them. Peer provenance does not replace live RPC registration after a hub bounce; empty machine RPC registries still need the keepalive/rpc-register heal in this PR.

Refuse to greenlight

  • "Zero manual restarts" for first gen / ghost registry / true legacy
  • Upgrade eligibility from advertised capabilities alone
  • Auto-exec from unverified peer chips
  • Claiming peer provenance fixes empty RpcRegistry

@cloudflare-workers-and-pages

cloudflare-workers-and-pages Bot commented Jul 20, 2026 •

Copy link
Copy Markdown

Deploying hapi with  Cloudflare Pages  Cloudflare Pages

Latest commit: 9956836
Status: ✅  Deploy successful!
Preview URL: https://a2ec4fdc.hapi-bqd.pages.dev
Branch Preview URL: https://fix-hub-runner-version-gover.hapi-bqd.pages.dev

View logs

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Findings

  • [Major] Guard self-upgrade by the runner-self-upgrade capability — upgradeMachineRunner calls runner-self-upgrade for any machine missing the required Cursor capability, but runners old enough to be skewed often do not have this newly-added RPC registered. The automatic attempt then fails every cooldown and the UI Upgrade button returns a 502 instead of offering a viable path. Evidence: hub/src/sync/syncEngine.ts:587.
    Suggested fix:
    import { MACHINE_CAPABILITIES } from '@hapi/protocol/runnerCapabilities'
    
    const capabilities = machine.metadata?.capabilities ?? []
    if (!capabilities.includes(MACHINE_CAPABILITIES.RunnerSelfUpgrade)) {
        return {
            type: 'error',
            message: 'Runner does not support self-upgrade; upgrade the CLI manually and restart the runner',
            code: 'upgrade_unavailable',
        }
    }
  • [Major] Catch missing bun so npm fallback can run — installFromNpm intends to try bun add -g and then fall back to npm install -g, but runCommand('bun', ...) can throw before returning when bun is not on PATH. That makes npm-channel fleet upgrades fail on npm/npx installs that have npm available but not bun. Evidence: cli/src/upgrade/selfUpgrade.ts:64.
    Suggested fix:
    async function runCommand(command: string, args: string[]): Promise<{ ok: boolean; output: string }> {
        try {
            const proc = Bun.spawn([command, ...args], {
                stdout: 'pipe',
                stderr: 'pipe',
                env: process.env,
            })
            const [stdout, stderr, exitCode] = await Promise.all([
                new Response(proc.stdout).text(),
                new Response(proc.stderr).text(),
                proc.exited,
            ])
            return { ok: exitCode === 0, output: `${stdout}\n${stderr}`.trim() }
        } catch (error) {
            return { ok: false, output: error instanceof Error ? error.message : String(error) }
        }
    }

Questions

  • None.

Summary

  • Review mode: initial
  • The new skew soft-fail path is reasonable, but fleet upgrade currently fails for the exact older runners that lack the new self-upgrade RPC, and npm fallback does not actually fall back if bun is missing.

Testing

  • Not run (automation; static review only).

HAPI Bot

Comment thread hub/src/sync/syncEngine.ts Outdated
Comment thread cli/src/upgrade/selfUpgrade.ts Outdated

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Findings

  • [Major] Preserve runner argv during self-upgrade relaunch — scheduleRunnerRelaunch always starts the replacement as runner start-sync, so a runner originally started with --workspace-root ... comes back unscoped after Upgrade. That drops the workspace-root restrictions documented for browse/spawn and regresses the same handoff invariant the runner heartbeat already preserves via startedWithArgv. Evidence: cli/src/upgrade/selfUpgrade.ts:159.
    Suggested fix:
    import { readRunnerState } from '@/persistence'
    
    async function scheduleRunnerRelaunch(cliExecutable?: string): Promise<void> {
        const state = await readRunnerState()
        const args = Array.isArray(state?.startedWithArgv) && state.startedWithArgv[0] === 'runner'
            ? state.startedWithArgv
            : ['runner', 'start-sync']
        const env = {
            ...process.env,
            HAPI_RUNNER_HANDOFF_FROM_PID: String(process.pid),
        }
        const child = cliExecutable
            ? spawn(cliExecutable, args, { detached: true, stdio: 'ignore', env })
            : spawnHappyCLI(args, { detached: true, stdio: 'ignore', env })
        child.unref()
        setTimeout(() => process.exit(0), 250)
    }
    
    await scheduleRunnerRelaunch(installedExecutable)
  • [Major] Do not publish a 0.0.0 npm upgrade offer — when the hub is not running from a monorepo and the caller does not pass targetVersion, resolveUpgradeOffer falls back to 0.0.0. The real callers added in this PR omit targetVersion, so documented npm installs can ask runners to install @twsxtd/hapi@0.0.0, making every fleet-upgrade attempt fail. Evidence: hub/src/upgrade/resolveUpgradeOffer.ts:81.
    Suggested fix:
    // cli/src/commands/hub.ts
    import packageJson from '../../package.json'
    
    const hub = await startHub({
        args: context.commandArgs,
        cliVersion: packageJson.version,
    })
    
    // hub/src/startHub.ts
    export interface StartHubOptions {
        args?: string[]
        cliVersion?: string
    }
    
    const resolveCurrentUpgradeOffer = () => resolveUpgradeOffer({
        hubPackageRoot: defaultHubPackageRoot(),
        execPath: process.execPath,
        targetVersion: options.cliVersion,
    })

Questions

  • None.

Summary

  • Review mode: follow-up after new commits
  • The two previous bot findings were addressed, but the current self-upgrade path still has two Major regressions: replacement runners lose their original launch flags, and packaged hubs can emit an unusable npm target version.

Testing

  • Not run (automation; static review only). git diff --check origin/main...HEAD passed.

HAPI Bot

Comment thread cli/src/upgrade/selfUpgrade.ts Outdated
Comment thread hub/src/upgrade/resolveUpgradeOffer.ts Outdated

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Findings

  • [Major] Do not skip self-upgrade when capabilities are still behind — the runner returns already-current as soon as localVersion === offer.targetVersion, but the hub starts fleet upgrade because the machine is missing offer.targetCapabilities. Same-version source/soup builds, or any missed version bump, will be reported as successfully handled while the runner remains skewed; the hub also treats already-current as success, so the banner/auto-upgrade can churn without fixing the missing RPC. Evidence: cli/src/upgrade/selfUpgrade.ts:37, related handling hub/src/sync/syncEngine.ts:601.
    Suggested fix:
    import { CURRENT_MACHINE_CAPABILITIES } from "@hapi/protocol/runnerCapabilities"
    
    function hasTargetCapabilities(offer: HubUpgradeOffer, localCapabilities: readonly string[]): boolean {
        const local = new Set(localCapabilities)
        return offer.targetCapabilities.every((cap) => local.has(cap))
    }
    
    export function shouldApplyUpgradeOffer(
        offer: HubUpgradeOffer,
        localVersion: string,
        localCapabilities: readonly string[] = CURRENT_MACHINE_CAPABILITIES,
    ): ApplyDecision {
        // existing channel validation...
        if (localVersion === offer.targetVersion && hasTargetCapabilities(offer, localCapabilities)) {
            return { apply: false, reason: "already-current" }
        }
        return { apply: true, reason: "upgrade" }
    }

Questions

  • None.

Summary

  • Review mode: follow-up after new commits
  • The previous bot findings are addressed. One remaining Major issue: self-upgrade is version-only, so a runner can stay capability-skewed even after the hub asks it to upgrade.

Testing

  • Not run (automation; static review only).

HAPI Bot

Comment thread cli/src/upgrade/selfUpgrade.ts Outdated

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Findings

  • None.

Questions

  • None.

Summary

  • Review mode: follow-up after new commits
  • Previous HAPI Bot finding is addressed in the latest head. Full current diff reviewed; no new high-confidence issues found. Residual risk: mechanical validation could not be completed in this runner because bun is unavailable.

Testing

  • git diff --check b74a11ecc39451b1ae92cbfc7fb98744227a204a...368a364266e6c860c70ef33ce03a6c162020ed42 passed.
  • bun typecheck not run: bun command not found.

HAPI Bot

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Findings

  • None.

Questions

  • None.

Summary
Review mode: follow-up after new commits
Full current diff reviewed; no high-confidence issues found. Residual risk: runner self-upgrade install/relaunch paths were reviewed statically but not exercised in this runner.

Testing

  • git diff --check origin/main...HEAD passed.
  • Not run: bun typecheck / bun run test (bun not found in runner).

HAPI Bot

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Findings

  • [Major] Version-drifted runners can disappear from the upgrade UI — the hub now treats pure version drift and missing target capabilities as upgrade-required, but the banner only checks isMachineCapabilitySkewed, whose required set is just cursor-chat-store-status. A runner that already has that capability but is behind on version or lacks runner-self-upgrade can hit the auto-upgrade path, fail with upgrade_unavailable, and still never show the banner or UPDATE REQUIRED affordance. Evidence web/src/components/RunnerVersionSkewBanner.tsx:27, hub/src/sync/syncEngine.ts:482, shared/src/upgradeChannel.ts:106.
    Suggested fix:
    import { machineTrailsUpgradeOffer, type HubUpgradeOffer } from '@hapi/protocol/upgradeChannel'
    
    export function listSkewedMachines(machines: Machine[], offer: HubUpgradeOffer | null): Machine[] {
        return machines.filter((machine) => (
            machine.active
            && (
                isMachineCapabilitySkewed(machine.metadata?.capabilities)
                || (offer
                    ? machineTrailsUpgradeOffer(offer, machine.metadata?.happyCliVersion, machine.metadata?.capabilities)
                    : false)
            )
        ))
    }
    Wire offer from /api/upgrade/offer, and use the same predicate for the machine selector’s UPDATE REQUIRED label.

Questions

  • None.

Summary
Review mode: follow-up after new commits
Found one Major issue: frontend skew detection does not match the backend fleet-upgrade predicate, so some out-of-date runners can lose the visible recovery path.

Testing

  • git diff --check origin/main...HEAD passed.
  • Not run: bun typecheck / bun run test (review automation; PR code not executed).

HAPI Bot

Comment thread web/src/components/RunnerVersionSkewBanner.tsx

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Findings

  • [Major] Fleet policy updates fail from standalone web origins — the new policy endpoint is PUT, and the web client calls it with JSON, but the hub CORS middleware still allows only GET, POST, PATCH, DELETE, OPTIONS. For documented standalone hosting, the browser preflight for PUT /api/upgrade/policy is rejected before auth, so users cannot change the fleet-upgrade policy from the remote web app. Evidence hub/src/web/routes/upgrade.ts:31, context hub/src/web/server.ts:234.
    Suggested fix:

    const corsMiddleware = cors({
        origin: corsOriginOption,
        allowMethods: ['GET', 'POST', 'PUT', 'PATCH', 'DELETE', 'OPTIONS'],
        allowHeaders: ['authorization', 'content-type']
    })
  • [Minor] Machine selector still uses the old skew predicate — the banner now uses machineTrailsUpgradeOffer plus the policy, but the new-session selector labels UPDATE REQUIRED only for missing required capabilities and ignores silent. Pure version drift will be selectable without the label, while silent can still show an alert label for capability drift. Evidence web/src/components/NewSession/MachineSelector.tsx:17.
    Suggested fix:

    import { machineTrailsUpgradeOffer, type FleetUpgradePolicy, type HubUpgradeOffer } from '@hapi/protocol/upgradeChannel'
    
    function getMachineOptionLabel(
        machine: Machine,
        offer: HubUpgradeOffer | null,
        policy: FleetUpgradePolicy,
    ): string {
        const skew = policy !== 'silent'
            && offer
            && machine.active
            && machineTrailsUpgradeOffer(offer, machine.metadata?.happyCliVersion, machine.metadata?.capabilities)
            ? ' · UPDATE REQUIRED'
            : ''
        return `${title}${platform}${version}${skew}`
    }

Summary
Review mode: follow-up after new commits
The prior banner predicate issue is fixed for RunnerVersionSkewBanner, but the same mismatch remains in MachineSelector. The new fleet policy route also needs CORS coverage for the documented standalone-web deployment path.

Testing

  • Not run (automation; PR code not executed).

HAPI Bot

Comment thread hub/src/web/routes/upgrade.ts Outdated
Comment thread web/src/components/NewSession/MachineSelector.tsx

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Findings

  • [Major] Fleet policy updates fail from standalone web origins — the new policy endpoint is PUT, and the web client sends JSON to it, but the hub CORS middleware still allows only GET, POST, PATCH, DELETE, OPTIONS. For documented standalone hosting, the browser preflight for PUT /api/upgrade/policy is rejected before auth, so users cannot change the fleet-upgrade policy from the remote web app. Evidence hub/src/web/routes/upgrade.ts:31, context hub/src/web/server.ts:234.
    Suggested fix:

    const corsMiddleware = cors({
        origin: corsOriginOption,
        allowMethods: ['GET', 'POST', 'PUT', 'PATCH', 'DELETE', 'OPTIONS'],
        allowHeaders: ['authorization', 'content-type']
    })
  • [Minor] Machine selector still uses the old skew predicate — the banner and auto-upgrade path now use machineTrailsUpgradeOffer plus the fleet policy, but the new-session selector labels UPDATE REQUIRED only for missing required capabilities and ignores silent. Pure version drift will be selectable without the label, while silent can still show an alert label for capability drift. Evidence web/src/components/NewSession/MachineSelector.tsx:17.
    Suggested fix:

    import { machineTrailsUpgradeOffer, type FleetUpgradePolicy, type HubUpgradeOffer } from '@hapi/protocol/upgradeChannel'
    
    function getMachineOptionLabel(
        machine: Machine,
        offer: HubUpgradeOffer | null,
        policy: FleetUpgradePolicy,
    ): string {
        const skew = policy !== 'silent'
            && offer
            && machine.active
            && machineTrailsUpgradeOffer(offer, machine.metadata?.happyCliVersion, machine.metadata?.capabilities)
            ? ' · UPDATE REQUIRED'
            : ''
        return `${title}${platform}${version}${skew}`
    }

Summary
Review mode: follow-up after new commits
The latest follow-up moved runner-management settings copy, but these two issues remain on current head. The prior banner predicate issue is fixed for RunnerVersionSkewBanner; the mismatch remains in MachineSelector.

Testing

  • Not run (automation; PR code not executed).

HAPI Bot

Comment thread hub/src/web/routes/upgrade.ts Outdated
Comment thread web/src/components/NewSession/MachineSelector.tsx

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Findings

  • [Major] Self-upgrade exits before runner cleanup — scheduleRunnerRelaunch schedules process.exit(0) after 250 ms, while the caller delays requestShutdown by 500 ms, so the runner bypasses the normal cleanup path that updates state, stops the control server, and releases/removes runner state/lock files. A failed replacement can leave the machine offline or stuck behind stale handoff state. Evidence cli/src/upgrade/selfUpgrade.ts:194, context cli/src/api/apiMachine.ts:380.
    Suggested fix:

    async function scheduleRunnerRelaunch(cliExecutable?: string): Promise<void> {
        // spawn replacement as today
        child.unref()
    }
    
    await scheduleRunnerRelaunch(installedExecutable)
    if (options.requestShutdown) {
        options.requestShutdown()
    } else {
        setTimeout(() => process.exit(0), 500)
    }
  • [Major] Fleet policy updates fail from standalone web origins — the new policy endpoint is PUT, and the web client sends JSON to it, but the hub CORS middleware still allows only GET, POST, PATCH, DELETE, OPTIONS. For documented standalone hosting, the browser preflight for PUT /api/upgrade/policy is rejected before auth, so users cannot change the fleet-upgrade policy from the remote web app. Evidence hub/src/web/routes/upgrade.ts:31, context hub/src/web/server.ts:234.
    Suggested fix:

    const corsMiddleware = cors({
        origin: corsOriginOption,
        allowMethods: ['GET', 'POST', 'PUT', 'PATCH', 'DELETE', 'OPTIONS'],
        allowHeaders: ['authorization', 'content-type']
    })
  • [Major] Machine re-registration can keep stale capabilities — mergeMachineRegistrationMetadata preserves every old metadata field when the incoming runner omits it. If a runner that does not advertise capabilities reconnects or downgrades against a row that previously had them, the hub can keep believing runner-self-upgrade / cursor-chat-store-status are present and then hide skew or route RPCs to handlers that are not registered. Evidence shared/src/machineRegistration.ts:87.
    Suggested fix:

    export function mergeMachineRegistrationMetadata(existing: unknown, incoming: unknown): Record<string, unknown> {
        const current = asRecord(existing) ?? {}
        const next = asRecord(incoming) ?? {}
        const merged: Record<string, unknown> = { ...next }
        if (next.displayName === undefined && current.displayName !== undefined) {
            merged.displayName = current.displayName
        }
        return merged
    }
  • [Minor] Machine selector still uses the old skew predicate — the banner and auto-upgrade path now use machineTrailsUpgradeOffer plus the fleet policy, but the new-session selector labels UPDATE REQUIRED only for missing required capabilities and ignores silent. Pure version drift will be selectable without the label, while silent can still show an alert label for capability drift. Evidence web/src/components/NewSession/MachineSelector.tsx:17.
    Suggested fix:

    const skew = policy !== 'silent'
        && offer
        && machine.active
        && machineTrailsUpgradeOffer(offer, machine.metadata?.happyCliVersion, machine.metadata?.capabilities)
        ? ' · UPDATE REQUIRED'
        : ''

Questions

  • None.

Summary

  • Review mode: follow-up after new commits
  • The latest follow-up changes default fleet policy and moves runner-management settings, but the previous CORS and selector issues remain. I also found a self-upgrade shutdown race and a stale machine metadata merge problem in the current full diff.

Testing

  • Not run (review automation; PR code not executed).

HAPI Bot

Comment thread cli/src/upgrade/selfUpgrade.ts Outdated
Comment thread hub/src/web/routes/upgrade.ts Outdated
Comment thread shared/src/machineRegistration.ts Outdated
Comment thread web/src/components/NewSession/MachineSelector.tsx

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Findings

  • [Major] Self-upgrade exits before runner cleanup — scheduleRunnerRelaunch still schedules process.exit(0) after 250 ms, while the only caller delays requestShutdown by 500 ms. That bypasses the runner's normal shutdown path, so the old process can exit before state/lock cleanup and before requestShutdown gets a chance to run. Evidence cli/src/upgrade/selfUpgrade.ts:212, context cli/src/api/apiMachine.ts:438.
    Suggested fix:

    async function scheduleRunnerRelaunch(cliExecutable?: string): Promise<void> {
        // spawn replacement as today
        child.unref()
    }
    
    await scheduleRunnerRelaunch(installedExecutable)
    if (options.requestShutdown) {
        options.requestShutdown()
    } else {
        setTimeout(() => process.exit(0), 500)
    }
  • [Major] Fleet policy updates fail from standalone web origins — the new policy endpoint is PUT, and the web client sends JSON to it, but the hub CORS middleware still allows only GET, POST, PATCH, DELETE, OPTIONS. In the documented standalone web setup, browser preflight for PUT /api/upgrade/policy fails before auth, so operators cannot change this setting from the remote web app. Evidence hub/src/web/routes/upgrade.ts:31, context hub/src/web/server.ts:234.
    Suggested fix:

    const corsMiddleware = cors({
        origin: corsOriginOption,
        allowMethods: ['GET', 'POST', 'PUT', 'PATCH', 'DELETE', 'OPTIONS'],
        allowHeaders: ['authorization', 'content-type']
    })
  • [Major] Machine re-registration can preserve stale capabilities — the refresh path starts the merge from stored metadata, and machineRegistrationNeedsRefresh ignores cases where the incoming runner omits capabilities. A runner that reconnects/downgrades without advertising capabilities can keep old runner-self-upgrade / cursor-chat-store-status metadata, causing the hub to hide skew or attempt RPCs the runner does not actually register. Evidence shared/src/machineRegistration.ts:87.
    Suggested fix:

    if (!capabilitiesEqual(current.capabilities, next.capabilities)) {
        return true
    }
    
    export function mergeMachineRegistrationMetadata(existing: unknown, incoming: unknown): Record<string, unknown> {
        const current = asRecord(existing) ?? {}
        const next = asRecord(incoming) ?? {}
        const merged: Record<string, unknown> = { ...next }
        if (next.displayName === undefined && current.displayName !== undefined) {
            merged.displayName = current.displayName
        }
        return merged
    }
  • [Minor] Machine selector still uses the old skew predicate — the banner and upgrade flow use machineTrailsUpgradeOffer plus the fleet policy, but the new-session selector labels UPDATE REQUIRED only for missing required capabilities. Pure version drift is selectable without the label, while silent policy can still show the label for capability drift. Evidence web/src/components/NewSession/MachineSelector.tsx:17.
    Suggested fix:

    const skew = policy !== 'silent'
        && offer
        && machine.active
        && machineTrailsUpgradeOffer(offer, machine.metadata?.happyCliVersion, machine.metadata?.capabilities)
        ? ' · UPDATE REQUIRED'
        : ''

Questions

  • None.

Summary

  • Review mode: follow-up after new commits
  • The Windows artifact follow-up changes do not resolve the earlier shutdown, CORS, stale capability merge, or selector-label issues. I did not find additional high-confidence findings beyond those current-head problems.

Testing

  • Not run (review automation; PR code not executed).

HAPI Bot

Comment thread cli/src/upgrade/selfUpgrade.ts Outdated
Comment thread hub/src/web/routes/upgrade.ts Outdated
Comment thread shared/src/machineRegistration.ts Outdated
Comment thread web/src/components/NewSession/MachineSelector.tsx
heavygee pushed a commit to heavygee/hapi that referenced this pull request Jul 22, 2026
Ops call for tiann#1108 dogfood: hapi-runner-from-active ignores Upgrade
binaries and defaults HAPI_DISABLE_VERSION_HANDOFF=1. Rebuild auto-cleans
web/dist.prev churn; restart heals stale switch + lists effective WORKING.

Co-authored-by: Cursor <cursoragent@cursor.com>

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Findings

  • [Major] Self-upgrade exits before runner cleanup — scheduleRunnerRelaunch still schedules process.exit(0) after 250 ms, while the only caller delays requestShutdown by 500 ms. That bypasses the runner's normal shutdown path, so the old process can exit before state/lock cleanup and before requestShutdown gets a chance to run. Evidence cli/src/upgrade/selfUpgrade.ts:212, context cli/src/api/apiMachine.ts:380.
    Suggested fix:

    async function scheduleRunnerRelaunch(cliExecutable?: string): Promise<void> {
        // spawn replacement as today
        child.unref()
    }
    
    await scheduleRunnerRelaunch(installedExecutable)
    if (options.requestShutdown) {
        options.requestShutdown()
    } else {
        setTimeout(() => process.exit(0), 500)
    }
  • [Major] Fleet policy updates fail from standalone web origins — the new policy endpoint is PUT, and the web client sends JSON to it, but the hub CORS middleware still allows only GET, POST, PATCH, DELETE, OPTIONS. In the documented standalone web setup, browser preflight for PUT /api/upgrade/policy fails before auth, so operators cannot change this setting from the remote web app. Evidence hub/src/web/routes/upgrade.ts:31, context hub/src/web/server.ts:234.
    Suggested fix:

    const corsMiddleware = cors({
        origin: corsOriginOption,
        allowMethods: ['GET', 'POST', 'PUT', 'PATCH', 'DELETE', 'OPTIONS'],
        allowHeaders: ['authorization', 'content-type']
    })
  • [Major] Machine re-registration can preserve stale capabilities — the refresh path only compares capabilities when the incoming metadata includes them, then the merge starts from stored metadata. A runner that reconnects/downgrades without advertising capabilities can keep old runner-self-upgrade / cursor-chat-store-status metadata, causing the hub to hide skew or attempt RPCs the runner does not actually register. Evidence shared/src/machineRegistration.ts:87.
    Suggested fix:

    if (!capabilitiesEqual(current.capabilities, next.capabilities)) {
        return true
    }
    
    export function mergeMachineRegistrationMetadata(existing: unknown, incoming: unknown): Record<string, unknown> {
        const current = asRecord(existing) ?? {}
        const next = asRecord(incoming) ?? {}
        const merged: Record<string, unknown> = { ...next }
        if (next.displayName === undefined && current.displayName !== undefined) {
            merged.displayName = current.displayName
        }
        return merged
    }
  • [Minor] Machine selector still uses the old skew predicate — the banner and upgrade flow use machineTrailsUpgradeOffer plus the fleet policy, but the new-session selector labels UPDATE REQUIRED only for missing required capabilities. Pure version drift is selectable without the label, while silent policy can still show the label for capability drift. Evidence web/src/components/NewSession/MachineSelector.tsx:17.
    Suggested fix:

    const skew = policy !== 'silent'
        && offer
        && machine.active
        && machineTrailsUpgradeOffer(offer, machine.metadata?.happyCliVersion, machine.metadata?.capabilities)
        ? ' · UPDATE REQUIRED'
        : ''

Questions

  • None.

Summary
Review mode: follow-up after new commits
The latest follow-up adds versionHandoffDisabled handling and banner filtering, but the current head still contains the shutdown, CORS, stale capability merge, and selector-label issues above. I did not find additional high-confidence findings beyond these current-head problems.

Testing

  • Not run (review automation; PR code not executed).

HAPI Bot

Comment thread cli/src/upgrade/selfUpgrade.ts Outdated
Comment thread hub/src/web/routes/upgrade.ts Outdated
Comment thread shared/src/machineRegistration.ts Outdated
Comment thread web/src/components/NewSession/MachineSelector.tsx
@heavygee

Copy link
Copy Markdown
Collaborator Author

Review debt cleared — tip 5114dc30c

Addressed the repeated Codex Major cluster (duplicates collapsed to one fix each):

Theme Status Evidence
1. scheduleRunnerRelaunch hard process.exit(0) @250ms races requestShutdown @500ms Fixed Spawn only in scheduleRunnerRelaunch; applyRunnerSelfUpgrade calls requestShutdown and only hard-exits (500ms) when no shutdown hook is provided
2. CORS omits PUT for /api/upgrade/policy Fixed hub/src/web/server.ts allowMethods now includes PUT
3. machineRegistration preserves stale capabilities Fixed Merge starts from registering payload; omitted capabilities → []. needsRefresh always compares capabilities (omit = empty)
4. Banner predicate narrower than machineTrailsUpgradeOffer Already fixed in tip before this commit (d56577879 / earlier) — banner uses machineTrailsUpgradeOffer + versionHandoffDisabled exclude. Bot thread was stale vs tip.
Minor: MachineSelector capability-only skew Fixed Same offer/policy/handoff gates as banner (machineNeedsUpdateLabel)

Tests: machineRegistration, fleet upgrade, banner, MachineSelector unit; bun typecheck green.

Resolved all open review threads on this PR (were 17). Please re-run Codex if needed — expect Majors to clear.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Findings

  • No issues found at >=80% confidence in the current diff.

Questions

  • None.

Summary
Review mode: follow-up after new commits
The follow-up changes address the prior bot findings I checked: graceful runner shutdown, CORS for PUT /api/upgrade/policy, stale capability clearing on re-registration, and selector labeling via the offer/policy predicate. Residual risk is the broad fleet-upgrade surface across CLI, hub, shared contracts, and web UI.

Testing

  • Not run (review automation; PR code not executed).

HAPI Bot

@heavygee

Copy link
Copy Markdown
Collaborator Author

Rebase onto upstream/main (Release 0.23.4)

Meta fleet: was DIRTY/CONFLICTING (~47 behind). Rebased onto upstream/main @ 8eac26726.

  • New tip: 43b3e82873edb25e2201771f4c8f5ae555135882 (43b3e8287)
  • Conflict: cli/src/agent/sessionFactory.ts on first commit only — rerere prior resolution preserved CURRENT_MACHINE_CAPABILITIES + mtime fields alongside HAPI_SESSION_ID export from feat(cli): export HAPI_SESSION_ID into wrapped agent env (self-targeting) #1121.
  • Preserved Majors from prior green tip 5114dc30c: exit-race (no early process.exit in scheduleRunnerRelaunch), CORS PUT, stale-capability merge from registering payload, machineTrailsUpgradeOffer / handoff-disabled soup path.
  • Verification: bun typecheck OK; full cli+hub+web suites green in worktree; focused shared/hub upgrade + machineRegistration / selfUpgrade / runnerEnsure tests pass.

Still do not merge — wait on @tiann. Bot may re-review this tip.

@heavygee
heavygee force-pushed the fix/hub-runner-version-governance branch from 5114dc3 to 43b3e82 Compare July 24, 2026 09:05

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Findings

  • No issues found at >=80% confidence in the current diff.

Questions

  • None.

Summary
Review mode: follow-up after new commits
Reviewed the full PR diff against main plus the prior HAPI Bot review context. The follow-up head still looks consistent across runner capability advertisement, hub upgrade offer/policy routing, self-upgrade gating, Cursor reopen soft-fail behavior, and the web banner/settings wiring. Residual risk is the broad fleet-upgrade path because review automation did not execute PR code.

Testing

  • Not run (review automation; PR code not executed).

HAPI Bot

@heavygee

Copy link
Copy Markdown
Collaborator Author

Rebase onto upstream/main again

Was DIRTY/CONFLICTING (~27 behind). Rebased onto upstream/main @ d53e25700.

  • New tip: 27d6a5476a8024b844c9354a101436b9a6c1a2f1 (27d6a5476)
  • Conflicts resolved:
  • Follow-up commit: SettingsPageContent no longer takes title (moved to SettingsSection); metadata backfill tests expect normalized capabilities: []
  • Majors preserved (exit-race, CORS PUT, stale-capability merge, handoff/soup path)
  • Verify: bun typecheck OK; focused machineRegistration / upgrade / machines / selfUpgrade tests green

Still do not merge — wait on @tiann.

@heavygee
heavygee force-pushed the fix/hub-runner-version-governance branch from 43b3e82 to 27d6a54 Compare July 27, 2026 10:16

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Findings

  • [Major] Terminal session registration can mask a stale runner — getOrCreateMachine now refreshes capability/version metadata for every registration, but local terminal bootstraps also hit /cli/machines with runnerState: null while using the same machine id (cli/src/agent/sessionFactory.ts:167). If a newer local CLI starts a session while an older runner is still online, this path can replace the active runner row with the newer CLI's happyCliVersion/capabilities, so machineTrailsUpgradeOffer stops reporting skew and the banner/auto-upgrade path disappears even though the live runner socket is still old (hub/src/sync/syncEngine.ts:493). Evidence: hub/src/store/machines.ts:74.
    Suggested fix:
    const isRunnerRegistration = runnerState !== null && runnerState !== undefined
    
    if (isRunnerRegistration && machineRegistrationNeedsRefresh(stored.metadata, metadata)) {
        const merged = mergeMachineRegistrationMetadata(stored.metadata, metadata)
        // existing updateMachineMetadata path
    }
    
    if (!isRunnerRegistration) {
        return stored
    }

Questions

  • None.

Summary
Review mode: follow-up after new commits
Found one Major issue in the latest full PR diff. The new runner-governance metadata needs to be owned by runner registration/heartbeat only; terminal session bootstrap should not be able to rewrite the active runner identity row. Residual risk remains in the broad self-upgrade/artifact path because review automation did not execute PR code.

Testing

  • Not run (review automation; PR code not executed).

HAPI Bot

Comment thread hub/src/store/machines.ts Outdated
@heavygee

Copy link
Copy Markdown
Collaborator Author

Follow-up: gate machine identity refresh on runner registration

Addresses Codex Major on getOrCreateMachine (terminal bootstrap masking stale runner).

  • Tip: ac16b35cb2f15ede3025ef4275623a40193ff7d8 (ac16b35cb)
  • Terminal runnerState: null no longer rewrites version/capabilities
  • Tests cover the mask regression

Still wait on @tiann — do not merge.

@heavygee
heavygee force-pushed the fix/hub-runner-version-governance branch from ac16b35 to a74be5b Compare July 27, 2026 19:00
@heavygee

Copy link
Copy Markdown
Collaborator Author

Addressed the latest Codex Major (refreshMachine() drop live RPC overlay on cold-cache Upgrade fallback) in 1d48f39b3:

  • MachineCache.refreshMachine now returns withLiveCapabilities(...) (cache still stores bare metadata)
  • Regression tests: machineCache.liveCapabilities.test.ts, fleetUpgrade.test.ts cold-cache path

Prior Major (terminal registration masking) remains fixed since 4585e758f / runnerState gate.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Findings

  • [Major] npm self-upgrade relaunches the old CLI — the npm branch installs the target package, but it leaves installedExecutable undefined, so scheduleRunnerRelaunch() falls back to spawnHappyCLI(). In compiled mode that helper resolves the currently running executable (process.argv[0] / process.execPath) and propagates it via HAPI_CLI_EXECUTABLE, so the RPC can report started and stop the runner while the replacement process is still the old generation. Evidence cli/src/upgrade/selfUpgrade.ts:244, related context cli/src/utils/spawnHappyCLI.ts:97.
    Suggested fix:
    async function installFromNpm(offer: HubUpgradeOffer): Promise<string> {
        // existing bun/npm install logic
        return process.platform === 'win32' ? 'hapi.exe' : 'hapi'
    }
    
    if (options.offer.channel === 'npm') {
        installedExecutable = await installFromNpm(options.offer)
    }

Questions

  • None.

Summary
Review mode: follow-up after new commits
Found one Major issue in the latest full PR diff. The prior bot finding about terminal registration masking stale runner metadata is not repeated; the current code gates that refresh to runner registration.

Testing

  • Not run (review automation; PR code not executed).

HAPI Bot

Comment thread cli/src/upgrade/selfUpgrade.ts Outdated
@heavygee

Copy link
Copy Markdown
Collaborator Author

Addressed Codex Major (npm self-upgrade relaunches old CLI) in tip:

  • After bun add -g / npm install -g, resolve hapi on PATH and pass that absolute path into scheduleRunnerRelaunch (same direct-spawn path as hub-artifact)
  • Avoids spawnHappyCLI locking HAPI_CLI_EXECUTABLE to the still-running compiled binary

Tests: cli/src/upgrade/selfUpgrade.test.ts (resolvePostNpmInstallExecutable)

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review mode: follow-up after new commits

Findings

  • [Major] Windows npm self-upgrade can relaunch a .cmd shim without shell handling. resolvePostNpmInstallExecutable() can return hapi.cmd, but scheduleRunnerRelaunch() spawns that path directly. This repo already uses shell: true for Windows command shims elsewhere; without it, the npm-channel upgrade path will fail on Windows. cli/src/upgrade/selfUpgrade.ts:110-119,227-232
    Suggested fix:
    const needsShell = process.platform === 'win32' && /\.(cmd|bat)$/i.test(cliExecutable)
    const child = spawn(cliExecutable, args, {
        detached: true,
        stdio: 'ignore',
        env,
        shell: needsShell,
        windowsHide: needsShell,
    })
  • [Major] versionHandoffDisabled hosts are filtered out of the skew banner entirely, so the new Restart escape hatch never appears for the exact machines that need it. The backend still treats restartMachineRunner() as the escape hatch for handoff-disabled/soup hosts, but listSkewedMachines() drops them before render, and MachineSelector applies the same filter. web/src/components/RunnerVersionSkewBanner.tsx:50-56, web/src/components/NewSession/MachineSelector.tsx:27-30
    Suggested fix:
    return machines.filter((machine) =>
        machine.active && machineTrailsUpgradeOffer(offer, machine.metadata?.happyCliVersion, machine.metadata?.capabilities)
    )
    // then disable only the Upgrade action when versionHandoffDisabled is true

Summary
Review mode: follow-up after new commits. I found 2 regressions in the new upgrade flow. Residual risk: no Windows runtime smoke test in the diff, and no UI test covering the handoff-disabled restart path.

Testing
Not run (automation)
HAPI Bot

Comment thread cli/src/upgrade/selfUpgrade.ts Outdated
Comment thread web/src/components/RunnerVersionSkewBanner.tsx
@heavygee

Copy link
Copy Markdown
Collaborator Author

Addressed the two latest Codex Majors:

  1. Windows .cmd relaunch — prefer hapi.exe on PATH; if only a .cmd/.bat shim remains, scheduleRunnerRelaunch uses shell: true + windowsHide.
  2. handoff-disabled banner — soup/rebuild-only hosts stay in the skew list; Upgrade is disabled, Restart is enabled as the escape hatch. MachineSelector labels them too.

@heavygee
heavygee force-pushed the fix/hub-runner-version-governance branch from 4a1b6d4 to 0d9eaf0 Compare August 8, 2026 07:02

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Findings

  • None.

Questions

  • None.

Summary

Review mode: follow-up after new commits

No high-confidence issues found in the full latest diff. The follow-up head rebases the previously reviewed fleet-governance stack onto the current base; the modified integration points were re-checked along with the complete PR diff. Residual risk remains in real cross-platform supervisor handoffs and on-demand hub-artifact compilation/download behavior, which were reviewed statically but not exercised locally in this automation run.

Testing

  • Local tests: Not run (automation; contributor code was not executed per the security constraint).
  • git diff --check against the PR base: passed.
  • GitHub checks at posting time: both Test / test jobs passed; Cloudflare Pages passed.

HAPI Bot

swear01 added a commit to swear01/hapi that referenced this pull request Aug 9, 2026
Maintained release on top of official v0.27.2 (upstream main 00b6e44).
Carries: provider profiles tiann#1091, notifications tiann#987, searchable share tiann#986,
Codex agent budget tiann#847, fleet runner version governance tiann#1108, session jobs
tiann#1424, steer tiann#906, project groups tiann#955, and the maintained distribution
overlay. Drops tiann#1092 (workspace management) per operator decision.
Includes the desktop launcher packaging (tiann#769) gated behind the maintained
release flow.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Findings

  • [Minor] Coalesce upgrades per machine before invoking the runner - under auto, maybeFleetUpgradeMachine() starts this method asynchronously, while the still-visible banner and another browser can call the same public method concurrently. Both requests independently prepare an artifact and send runner-self-upgrade; the CLI's process-local gate rejects the second as "Runner upgrade already in progress", which the hub converts into upgrade_failed and may toast as a real fleet failure. Evidence: hub/src/sync/syncEngine.ts:1045

    Suggested fix:

    const existing = this.fleetUpgradeInFlight.get(machineId)
    if (existing) return await existing
    
    const task = this.upgradeMachineRunnerUnlocked(machineId, namespace)
    this.fleetUpgradeInFlight.set(machineId, task)
    try {
        return await task
    } finally {
        if (this.fleetUpgradeInFlight.get(machineId) === task) {
            this.fleetUpgradeInFlight.delete(machineId)
        }
    }

Questions

  • None.

Summary

Review mode: follow-up after new commits

One Minor issue found in the full latest diff. Without hub-side per-machine coalescing, overlapping automatic and manual upgrade requests can produce duplicate work and false failure reporting. Residual risk remains in real cross-platform supervisor handoffs and artifact installation paths, which were reviewed statically only.

Testing

  • Not run (automation; contributor code was not executed per the security constraint).
  • git diff --check against the PR base: passed.
  • GitHub checks at posting time: both test jobs passed; Cloudflare Pages passed.
  • Add a regression test that overlaps an auto-triggered attempt with a manual upgradeMachineRunner() call and asserts one artifact preparation/RPC plus a shared result.

HAPI Bot

Comment thread hub/src/sync/syncEngine.ts Outdated

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Findings

  • None.

Questions

  • None.

Summary
Review mode: follow-up after new commits

No issues found in the latest full diff. The prior Minor finding is resolved at the current head: concurrent automatic and manual runner upgrades now share one per-machine in-flight operation, with guarded cleanup after settlement.

Residual risk remains in cross-platform runner handoff, supervisor restart, and artifact installation behavior, which was reviewed statically only. The PR test plan also leaves operator dogfood incomplete.

Testing

  • Not run (automation; contributor code was not executed per the security constraint).
  • git diff --check against the PR base: passed.
  • GitHub checks at posting time: both test jobs passed; Cloudflare Pages passed; this review job was still in progress.
  • Coverage gap: the added coalescing regression overlaps two direct upgradeMachineRunner() calls; an end-to-end auto-heartbeat plus manual banner overlap remains untested.

HAPI Bot

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Findings

  • None.

Questions

  • None.

Summary
Review mode: follow-up after new commits

No issues found in the latest full diff. The new head addresses the ghost-RPC state by re-registering machine handlers on keepalive and refusing self-upgrade until the handler is present in the live registry.

Residual risk remains in cross-platform runner handoff, supervisor restart, and artifact build/install behavior, which was reviewed statically only. The PR test plan also leaves operator dogfood incomplete.

Testing

  • Not run (automation; contributor code was not executed per the security constraint).
  • git diff --check against the PR base: passed.
  • GitHub checks at posting time: both test jobs passed; Cloudflare Pages passed; this review job was still in progress.
  • Coverage gap: no direct integration test exercises ApiMachineClient keepalive re-registration through the hub before machine-alive; the added tests cover the live-registry gate, not the end-to-end healing path.

HAPI Bot

swear01 added a commit to swear01/hapi that referenced this pull request Aug 10, 2026
heavygee added a commit to heavygee/hapi that referenced this pull request Aug 10, 2026
Ghost RpcRegistry vs advertised caps, chicken/egg first restart,
npm vs artifact channels, soup landmines, and kill-criteria before
claiming smooth fleet deploy. Cross-link from reenroll runbook.

Co-authored-by: Cursor <cursoragent@cursor.com>
Debian and others added 4 commits August 11, 2026 12:03
Hub↔runner protocol drift was reported as missing Cursor chat data when
cursor-chat-store-status was unregistered. Soft-fail reopen on probe errors,
advertise required machine capabilities, surface an unmissable upgrade banner,
and stop-runner when a newer CLI binary is already on disk.

Fixes #1084

Co-authored-by: Cursor <cursoragent@cursor.com>
Compact the out-of-date banner (minimize + 1h snooze + per-host Restart)
so it no longer blocks the session list. Auto stop-runner on skew stays
opt-in via HAPI_AUTO_UPGRADE_RUNNERS / autoUpgradeRunners (default off).

Co-authored-by: Cursor <cursoragent@cursor.com>
QuotaExceededError from setItem aborted minimize before React state
updated, leaving the banner stuck over the session list. Persist to
memory when storage fails; only enable Restart when a newer CLI is
already on disk; clarify opt-in is stop-runner only, not package push.

Co-authored-by: Cursor <cursoragent@cursor.com>
…tart

CLI version handoff already reloads the runner when the on-disk binary
mtime changes. Hub-driven stop-runner on skew duplicated that. Keep the
skew banner and manual Restart only as a stuck/disabled-handoff escape.

Co-authored-by: Cursor <cursoragent@cursor.com>
@heavygee
heavygee force-pushed the fix/hub-runner-version-governance branch from fee9084 to bf12041 Compare August 11, 2026 12:52

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Findings

  • [Major] Ordinary terminal sessions can publish runner-only capabilities, hiding an actually outdated runner. Evidence: cli/src/agent/sessionFactory.ts:59.
  • [Major] The new Restart action only sends stop-runner, so an unsupervised runner can remain offline. Evidence: hub/src/sync/rpcGateway.ts:283.

Questions

  • None.

Summary
Review mode: follow-up after new commits

Two major issues found in the latest full PR diff. The previous HAPI Bot review was used only as context because its reviewed commit belongs to the earlier mirror history.

Testing

  • Not run (automation; contributor code was not executed per the security constraint).
  • git diff --check against the PR base: passed.
  • GitHub checks at posting time: both test jobs passed; Cloudflare Pages passed; this review job was still in progress.
  • Coverage gaps: no test proves terminal bootstrap cannot advertise machine-RPC capabilities for an older live runner; no integration test proves Restart produces a replacement PID and keeps the machine online.

HAPI Bot

Comment thread cli/src/agent/sessionFactory.ts
Comment thread hub/src/sync/rpcGateway.ts
Address #1108 bot Majors on the thin tip: terminal/lazy bootstraps no
longer merge CURRENT_MACHINE_CAPABILITIES into the machine row (only
asRunner registration does). Banner Restart refuses unsupervised hosts
so stop-runner cannot leave a detached laptop offline; supervised
runners advertise supervisedRestart via HAPI_RUNNER_SUPERVISED=1.

Co-authored-by: Cursor <cursoragent@cursor.com>

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Findings

  • [Major] Runner governance metadata is sticky across rollback/reconfiguration — newly advertised capabilities, mtimes, and supervisedRestart are merged over the stored machine row; absent values never clear. An older or newly unsupervised runner can therefore remain recorded as current/supervised, hiding skew and potentially making a later Restart stop an unsupervised runner. Evidence: cli/src/agent/sessionFactory.ts:72, related merge context hub/src/store/machines.ts:47.
    Suggested fix:
    const runnerOwnedKeys = ['capabilities', 'startedCliMtimeMs', 'installedCliMtimeMs', 'supervisedRestart'] as const
    const mergeBase = runnerState == null ? stored.metadata : omitKeys(stored.metadata, runnerOwnedKeys)
    const merged = mergeMachineMetadata(mergeBase, metadata)
    // Current runners should also send supervisedRestart: process.env.HAPI_RUNNER_SUPERVISED === '1'
  • [Minor] Canonical supervisor setups never set the new supervision marker — the installation text promises a per-host Restart escape hatch, but runtime requires HAPI_RUNNER_SUPERVISED=1; that variable is not present in the systemd/pm2/launchd examples or CLI configuration docs. Evidence: docs/guide/installation.md:331; configuration: Not found in repo/docs.
    Suggested fix:
    [Service]
    Environment=HAPI_RUNNER_SUPERVISED=1
    HAPI_RUNNER_SUPERVISED=1 pm2 start "hapi runner start-sync" --name hapi-runner
  • [Minor] The machine selector adds a hard-coded English upgrade status, so the Chinese UI renders UPDATE REQUIRED. Evidence: web/src/components/NewSession/MachineSelector.tsx:19.
    Suggested fix:
    const skew = machine.active && isMachineCapabilitySkewed(machine.metadata?.capabilities)
        ? ` · ${updateRequiredLabel}`
        : ''
    // Pass t('runner.skew.updateRequired') and add en/zh-CN locale entries.

Questions

  • None.

Summary
Review mode: follow-up after new commits

The two previous Major findings were rechecked against the full latest diff. The terminal-bootstrap capability issue is addressed. The unsupervised-stop guard is present, but persisted runner metadata can still become stale and bypass that guard after rollback or supervision changes.

Testing

  • Not run (automation; contributor code was not executed per the security constraint).
  • git diff --check against the PR base: passed.
  • GitHub checks at review time: Cloudflare Pages passed; one test job passed; one test job remained in progress; this review job remained in progress.
  • Coverage gap: no store/integration test re-registers the same machine from supervised/current to unsupervised or legacy metadata and verifies runner-owned governance fields are cleared.

HAPI Bot

Comment thread cli/src/agent/sessionFactory.ts
Comment thread docs/guide/installation.md
Comment thread web/src/components/NewSession/MachineSelector.tsx Outdated
… label

Omit-means-clear on runner registration so rollback cannot leave
supervisedRestart/capabilities sticky; always advertise boolean
supervisedRestart from asRunner. Document HAPI_RUNNER_SUPERVISED=1
and localize MachineSelector UPDATE REQUIRED.

Co-authored-by: Cursor <cursoragent@cursor.com>
@heavygee

Copy link
Copy Markdown
Collaborator Author

Review threads addressed — tip 99568368f

Thread Fix
[Major] Sticky runner governance metadata Runner registration clears omitted ads (clearOmittedRunnerAds); asRunner always sends supervisedRestart boolean
[Minor] SUPERVISED missing from docs systemd / pm2 / launchd examples + installation.md pointer
[Minor] MachineSelector English hard-code runner.skew.updateRequired en + zh-CN

Thin gate: 35 files / 6 commits vs upstream/main. Still do not merge — wait on @tiann + bot re-review.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Findings

  • None.

Summary
Review mode: follow-up after new commits

No issues found at the current head. The prior sticky runner-governance metadata, supervisor documentation, and machine-selector localization findings are resolved. Residual risk remains around the real supervisor cold-relaunch/reconnect path; coverage in this PR mocks the stop RPC rather than exercising systemd/pm2/launchd end to end.

Testing

  • Not run (automation; contributor code was not executed per the security constraint).
  • git diff --check against the PR base: passed.
  • GitHub checks at review time: both test jobs and Cloudflare Pages passed.
  • Coverage gap: no end-to-end supervised runner restart/handoff test; operator dogfood remains unchecked in the PR test plan.

HAPI Bot

@tiann
tiann merged commit 1cd4d11 into main Aug 11, 2026
7 checks passed
@tiann
tiann deleted the fix/hub-runner-version-governance branch August 11, 2026 14:24
heavygee added a commit to heavygee/hapi that referenced this pull request Aug 11, 2026
Absorb tiann#1108 runner ads and tiann#1480 steer. Keep tag/proof gates before
metadata refresh; runner registration still sends asRunner mtimes.

Co-authored-by: Cursor <cursoragent@cursor.com>
heavygee added a commit to heavygee/hapi that referenced this pull request Aug 11, 2026
Gate A after upstream merge of fix/hub-runner-version-governance
(squash 1cd4d11). Manifest DROPPED; Gate A' retro filed.

Co-authored-by: Cursor <cursoragent@cursor.com>
heavygee added a commit to heavygee/hapi that referenced this pull request Aug 11, 2026
Co-authored-by: Cursor <cursoragent@cursor.com>

tiann#1108 exit reflection: chip cache lags live classify until :00 hapi-meta-daily.
heavygee added a commit to heavygee/hapi that referenced this pull request Aug 11, 2026
Incident plan: latching 🛑 needs_operator chip (no peer ping) plus
pinned fat upgrade SHAs for estate soup. Playback pending confirm.

Co-authored-by: Cursor <cursoragent@cursor.com>
heavygee added a commit to heavygee/hapi that referenced this pull request Aug 12, 2026
The driver/fleet-runner-upgrade comment cited tiann#1108, so mw_manifest_pr_layer_active
kept Gate A dirty after thin merge cleanup. Reword to #122-only scope.

Co-authored-by: Cursor <cursoragent@cursor.com>
RiriAgent added a commit to mouriya-s-lab/hapi that referenced this pull request Aug 13, 2026
Adopt upstream machine capability redesign (tiann#1108): MachineMetadata.capabilities
is now the runner's RPC capability id array (consumed by isMachineCapabilitySkewed
and cliBinaryUpdatedOnDisk in RunnerVersionSkewBanner/MachineSelector); fork's OMP
flag moves out of capabilities to a top-level ompAvailable field. Update all fork
consumers: sessionFactory, run.ts, apiMachine, machineCache, omp-host-integration
routes, NewSession, OmpProviderSettingsRow, and their tests.

Keep both sides' additive changes: upstream share-transfer/search imports,
isTranscriptEcho, fleet-version mtime/supervisedRestart metadata, restart-runner
route, NotifySummaryText/SpeakSummaryButton, SteerQueuedMessageResponse and fork's
multi-user settings routes, cc-switch providers route, ompInputMode, ProbeAgentSkills
response, formatUsageSnapshotLabel, session-summary-in-chat flag, scratchlist drawer
props, queue Steer button, and fork CI (session-scroll spec + integration env).

Locales: keep fork wording per prior resolution precedent (2af3340, 3d8d695).

Co-Authored-By: Mouriya-Emma <85676458+Mouriya-Emma@users.noreply.github.com>
WarLikeLaux pushed a commit to WarLikeLaux/hapi that referenced this pull request Sep 27, 2026
…e, soft-fail reopen) (tiann#1108)

* fix(hub): govern runner capabilities so Cursor reopen soft-fails on skew

Hub↔runner protocol drift was reported as missing Cursor chat data when
cursor-chat-store-status was unregistered. Soft-fail reopen on probe errors,
advertise required machine capabilities, surface an unmissable upgrade banner,
and stop-runner when a newer CLI binary is already on disk.

Fixes tiann#1084

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web,hub): make runner skew banner dismissible; gate auto-upgrade

Compact the out-of-date banner (minimize + 1h snooze + per-host Restart)
so it no longer blocks the session list. Auto stop-runner on skew stays
opt-in via HAPI_AUTO_UPGRADE_RUNNERS / autoUpgradeRunners (default off).

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web): tolerate full sessionStorage on skew banner minimize

QuotaExceededError from setItem aborted minimize before React state
updated, leaving the banner stuck over the session list. Persist to
memory when storage fails; only enable Restart when a newer CLI is
already on disk; clarify opt-in is stop-runner only, not package push.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(hub): drop redundant autoUpgradeRunners; runners already self-restart

CLI version handoff already reloads the runner when the on-disk binary
mtime changes. Hub-driven stop-runner on skew duplicated that. Keep the
skew banner and manual Restart only as a stuck/disabled-handoff escape.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(cli,hub,web): runner-only caps ads; gate Restart on supervisor

Address tiann#1108 bot Majors on the thin tip: terminal/lazy bootstraps no
longer merge CURRENT_MACHINE_CAPABILITIES into the machine row (only
asRunner registration does). Banner Restart refuses unsupervised hosts
so stop-runner cannot leave a detached laptop offline; supervised
runners advertise supervisedRestart via HAPI_RUNNER_SUPERVISED=1.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(hub,cli,web): clear sticky runner ads; docs SUPERVISED; i18n skew label

Omit-means-clear on runner registration so rollback cannot leave
supervisedRestart/capabilities sticky; always advertise boolean
supervisedRestart from asRunner. Document HAPI_RUNNER_SUPERVISED=1
and localize MachineSelector UPDATE REQUIRED.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Debian <heavygee@oos-linux.in.lockhouse>
Co-authored-by: Cursor <cursoragent@cursor.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Fleet runner version governance: skew banner, self-upgrade, soft-fail Cursor reopen

2 participants