Skip to content

Measure how long people stay cut off when the server is replaced - #8

Open
namespaceMarcello wants to merge 3 commits into
basecamp:mainfrom
namespaceMarcello:cable-reconnect
Open

namespaceMarcello wants to merge 3 commits into
basecamp:mainfrom
namespaceMarcello:cable-reconnect

Conversation

@namespaceMarcello

@namespaceMarcello namespaceMarcello commented Oct 10, 2026 •

Copy link
Copy Markdown
Contributor

loadgen cable holds connections open and measures fan-out. Nothing here measures what happens to those connections when the server is restarted or replaced: how long each person stays cut off, and whether the clients come back spread out or all at once.

What it adds. loadgen reconnect --base URL --room ID --sessions FILE --clients N:

  • N people from --sessions subscribe as a room page does. When they are all confirmed the command prints PHASE ready on stderr, and the caller restarts or redeploys the server; the command doesn't touch the server itself.
  • Each client then gets back in as --policy says:
    • actioncable, the default, replays the browser: ConnectionMonitor and Connection of @rails/actioncable. No attempt when the socket closes, but at the next turn of the poll timer (every 6 to 12 s), never within 6 s of the close, then every 6 × 1.15ⁿ s with up to 15% of jitter; a socket that opened and has been silent for 6 s at a poll is closed and reopened 500 ms later, one still connecting is left alone. Each client's timer starts where one that has been running for hours would be: clients that all connected a few seconds earlier would otherwise poll in step.
    • herd retries every --retry-ms from the close on, with no wait: what the monitor exists to prevent, and what a freshly started server can take.
    • early is the monitor plus a retry of the page's own, --early-min-ms to --early-max-ms after the close.
  • A client is back when its subscriptions are confirmed again and its room refresh request has answered 200, the request refresh_room_controller.js makes when HeartbeatChannel is confirmed. A drop while it is getting back in voids both: the socket that gets in asks again.
  • The result has each person's time from socket close to being back (min, p10, p50, p90, p99, max), the first attempt, attempts made and failed, connections the monitor reopened, arrivals per second, and when /up stopped and resumed answering. --trace FILE writes one line per client.
  • It fails closed: a refresh that isn't a 200 leaves that client not back, everyone_back says whether all of them returned, and the command exits with an error in three cases: a client didn't subscribe, a socket closed before PHASE ready (in both the phase is not printed), or no socket closed after it.

How the replay was checked. Against headless Chrome 155 on the Rails implementation, on one machine (a laptop, Docker): 16 windows on a room over three docker restarts were back after 11.3 s at the median (6.4 to 15.3 s), one attempt each; the replay with 1,000 clients gave 10.8 s (first attempts from 6.0 to 15.5 s). It is a replay and not a browser, which the README now says. It was written to measure basecamp/once-campfire#358, whose numbers come from the first commit here; the second moves the monitor's rules into a struct of their own so they can be tested, and makes a refused refresh count. A 200-client run of the rebuilt binary behaves the same: first attempt at 6.06 s, median 10.0 s, everyone back. The third answers the review: an incomplete setup stops the command, a drop while getting back in voids the refresh, herd makes its first attempt at the close, and first_socket_open_secs is the first socket to open. Its binary was tried with 20 clients under each policy, and under herd also with the app dropping the connections itself: everyone back each time.

Tests. Fourteen new ones. Five on the monitor and the summaries: the monitor leaves a fresh close alone for six seconds, a failed attempt isn't another disconnect, a welcome starts the count again, the poll intervals stay in the monitor's bounds, the summaries. Seven with one person against a local server: one the server drops gets back in, resubscribes and refreshes with their own cookie; one whose refresh is answered 500 is not back; herd tries again at the close; a refresh answered before a drop, and one still waiting at a drop, are asked again; the first socket to open is kept when a later one gets in; a close before PHASE ready stops the client. Two through the command itself: it fails when one of two clients is turned away, and when a socket closes before PHASE ready. Each was seen failing with one thing broken at a time.

bin/check's commands pass: cargo fmt --check, cargo clippy --all-targets -- -D warnings, cargo test (34), the five Ruby tests and the four node --checks.

Written by Claude Code (AI assistant) on behalf of @namespaceMarcello, who directs this work.

🤖 Generated with Claude Code

namespaceMarcello and others added 2 commits October 10, 2026 15:17
N people hold a room page's cable connection; the caller replaces the server at PHASE ready, and each client gets back in as the browser does: a replay of Action Cable's ConnectionMonitor (no retry before 6 s, then at the turn of its poll timer), or all at once (herd), or with an earlier retry of the page's own (early). Reports each person's blackout, from the socket closing to the subscriptions confirmed and the room refreshed, and probes /up to tell when the server itself was away.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017uAnU1dboBiJTZ23xmGpPe
…ent as back

The rules replayed from Action Cable's connection monitor move into a
struct, so that they can be tested without a socket: a fresh close is left
alone for six seconds, a failed attempt is not another disconnect, a
welcome starts the count again.

A client whose refresh request isn't answered with a 200 is no longer
back, the result says whether everyone returned, and the command fails
when no socket closed after PHASE ready.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017uAnU1dboBiJTZ23xmGpPe

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Readiness and recovery accounting can accept invalid runs and produce inaccurate measurements.

5 open findings
What changed in this PR

Adds a restart-recovery load profile that measures client blackout, reconnection behavior, and server availability.

Changes:

  • Adds Action Cable, herd, and early reconnection policies.
  • Reports recovery timings, attempts, health transitions, and optional traces.
  • Adds CLI integration, documentation, and Rust tests.

[!TIP]
If you aren't ready for review, convert to a draft PR.
Click "Convert to draft" or run gh pr ready --undo.
Click "Ready for review" or run gh pr ready to reengage.

File Description
README.md Documents the reconnect profile and policies.
loadgen/​src/​reconnect.rs Implements reconnect measurement, reporting, and tests.
loadgen/​src/​main.rs Registers the new subcommand.

🧠 Review effort: Balanced


💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread loadgen/src/reconnect.rs
Comment on lines +404 to +408
if closed {
if !ctx.armed.load(Ordering::Relaxed) {
rec.early_close = true;
}
rec.close.get_or_insert(now);

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in 80b9689. A client whose socket closes before PHASE ready now stops there instead of reconnecting, and the command fails with sockets closed before PHASE ready: N before printing the phase, so the caller doesn't replace the server for a run that can't be measured. Counting those closes and arming are one step under a mutex: a close is either counted before PHASE ready or it belongs to the restart. The early_close flag and closed_before_ready in the result are gone with it.

Tests: a_close_before_ready_is_not_a_blackout (the client doesn't get back in, and nobody counts as back) and a_socket_closing_before_ready_ends_the_run.

In the 130 results kept from this work closed_before_ready was 0 every time (one machine: a laptop, Docker), so the numbers quoted in basecamp/once-campfire#358 don't come from a run of this kind.

Written by Claude Code (AI assistant) on behalf of @namespaceMarcello, who directs this work.

Comment thread loadgen/src/reconnect.rs Outdated
Comment on lines +409 to +412
// A drop while getting back in voids what this socket had confirmed.
if rec.back.is_none() {
rec.confirmed = None;
}

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in 80b9689. A drop while getting back in now clears refreshed as well and aborts the request still in flight, so the socket that gets in asks for the room itself, as the page does (#channelConnected runs at every HeartbeatChannel connect).

Two tests against a local server, one with the first request answered 200 before the drop and one with it unanswered: a_refresh_answered_before_a_drop_is_asked_again and a_refresh_waiting_at_a_drop_is_asked_again. Both expect a second refresh request before the client is back; the second also expects the client to hang up on the first.

In the 88 traces kept from this work (78,412 clients; one machine: a laptop, Docker), 51 clients had a socket the monitor reopened while they were getting back in, all of them under early. In each of the 51 the refresh that counted was asked after the socket that got in had been welcomed, so no result so far rests on a stale one.

Written by Claude Code (AI assistant) on behalf of @namespaceMarcello, who directs this work.

Comment thread loadgen/src/reconnect.rs Outdated
monitor.closed(now);
if reopen_at.is_none() {
match ctx.policy {
Policy::Herd => reopen_at = Some(now + ctx.retry.mul_f64(rng.unit())),

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in 80b9689: herd now makes its first attempt at the close itself, and waits --retry-ms only after an attempt fails, which is what the module comment and the README say. Test: herd_tries_again_at_the_close (first_attempt equals close).

The runs made so far with herd had the random wait. None of their numbers is quoted here or in basecamp/once-campfire#358.

Written by Claude Code (AI assistant) on behalf of @namespaceMarcello, who directs this work.

Comment thread loadgen/src/reconnect.rs
while ctx.ready.load(Ordering::Relaxed) + ctx.failed.load(Ordering::Relaxed) < clients && Instant::now() < until {
tokio::time::sleep(Duration::from_millis(50)).await;
}
let ready = ctx.ready.load(Ordering::Relaxed);

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in 80b9689: when fewer than N clients are subscribed, because one failed or because the wait ran out, the command now fails with only R of N clients subscribed (F failed) instead of printing PHASE ready. ready and failed_before_ready are gone from the result, where they could only say N and 0. cable reports failed and carries on; here the phase line is what the caller acts on, so it stops.

Test: ready_takes_everyone_subscribed (two clients, one turned away). In the 130 results kept from this work failed_before_ready was 0 every time (one machine: a laptop, Docker).

Written by Claude Code (AI assistant) on behalf of @namespaceMarcello, who directs this work.

Comment thread loadgen/src/reconnect.rs Outdated
"server": {
"down_at_secs": down_at.map(secs),
"up_at_secs": up_at.map(secs),
"first_socket_open_secs": records.iter().filter_map(|r| r.opened).min().map(secs),

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in 80b9689: first_opened keeps each client's first socket to open after the close, and server.first_socket_open_secs reads it. opened stays the socket that got in, which upgrade_secs and welcome_secs are measured on. Test: the_first_socket_to_open_is_not_forgotten.

Written by Claude Code (AI assistant) on behalf of @namespaceMarcello, who directs this work.

From the review of the pull request:

- A client that didn't subscribe, or whose socket closed before PHASE
  ready, leaves no run to measure: the command now fails before printing
  the phase, so that the caller doesn't replace the server for nothing.
  Such a client used to reconnect, and could count as back without
  having been through the restart.
- A drop while getting back in voids the refresh too, answered or still
  in flight: the socket that gets in asks for the room itself, as the
  page does at every connect.
- herd makes its first attempt at the close, as documented. It waited a
  random part of --retry-ms first.
- server.first_socket_open_secs reads each client's first socket to
  open, where it read the one that got in.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YQodepLWf5cJkDSVkRMb2h
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants