Skip to content

Reconnect soon after the server drops the cable connection - #358

Open
namespaceMarcello wants to merge 2 commits into
basecamp:mainfrom
namespaceMarcello:reconnect-soon-after-the-cable-drops
Open

namespaceMarcello wants to merge 2 commits into
basecamp:mainfrom
namespaceMarcello:reconnect-soon-after-the-cable-drops

Conversation

@namespaceMarcello

@namespaceMarcello namespaceMarcello commented Oct 10, 2026 •

Copy link
Copy Markdown
Contributor

When the server drops a room page's cable connection, Action Cable's client doesn't try again right away. Unless the tab becomes visible or the browser comes back online in the meantime, its connection monitor never reopens a connection within six seconds of the close, and then waits for the next turn of its poll timer, so that the clients of a server that went away don't all come back at once (rails/rails#40229). A person is back after 6 to 17 seconds, half of them after about 10. From the fifth second the room is offline and the composer disabled, so everyone sees it.

The server is often there the whole time:

  • ONCE updates an app by starting the new container on the same volume, waiting for /up, and having kamal-proxy switch over (basecamp/once, deployWithVolume), and the proxy closes the WebSockets of the old container at once (Target#Drain: "Cancel any hijacked requests immediately"). The new server is already answering when the connections drop.
  • Campfire closes connections itself, with reconnect: true: for every member of a room that is deleted (Room::DestroyJob) and for a person removed from a room (Membership's after_destroy_commit).

This makes the room page try once, one to two seconds after the connection drops:

  • refresh_room_controller.js already keeps the page's online and offline state, and already nudges the monitor when the browser comes back online. Its disconnected callback now also schedules one consumer.connect(), between one and two seconds later, at a random moment so that the people in a room don't arrive in the same instant.
  • It doesn't when the server said not to reconnect: Action Cable has stopped its monitor then and passes willAttemptReconnect: false.
  • It is one attempt. connect() does nothing when the connection is already open or opening, and when the attempt fails Action Cable's monitor carries on exactly as it does today.

What stays as it is:

  • A plain docker restart: the server is away for longer than that (4.3 s on our machine), the early attempt finds nothing, and people are back when the monitor says, as today.
  • Pages other than a room, which don't have this controller.
  • A tab left open across the upgrade, which keeps the old script until it reloads.

What it costs: the people of a room arrive within a second instead of spread over ten. A server that has just started took about 200 reconnections a second on our machine (one reconnection is the handshake, its eight subscriptions and the refresh request: 15 SELECTs and 1 UPDATE on an empty cache, 16 to 26 ms of CPU). With 1,000 people in one room /up went unanswered for about 5 s while they came back, with 2,000 for 5 to 8 s; under today's schedule the same work is spread out, and with 1,000 people a reconnection takes 0.06 s. In the runs below, up to 500 people in a room were all back before the composer would have been disabled.

One limit of "one attempt": it is one per dropped connection. A server that accepted the early connection and dropped it again without saying not to reconnect would be tried again a second or two later, where the monitor would have waited six. A proxy with no backend refuses the upgrade instead, and a refused upgrade isn't a drop the page hears about.

Measurements

From our machine: a laptop with an AMD Ryzen 9 7940HX, Docker. The production image built from main at edbc779 and from this branch (Thruster, 3 Puma workers with 5 threads, 3 Resque workers, Redis inside) on 4 CPUs, the database on a Docker volume (ext4), plain HTTP, so no TLS handshakes are in these numbers.

In a browser. Headless Chrome 155, 16 windows on a room, three rounds, from the socket closing to the page being online again with its refresh answered; median, lowest and highest of the 48:

main this PR
the app drops the connections (User#reset_remote_connections) 10.7 s (6.4–15.8) 1.5 s (1.0–2.0)
windows that had the composer disabled 48 of 48 0 of 48
docker restart 11.3 s (6.4–15.3) 10.4 s (6.5–15.7)
windows that had the composer disabled 48 of 48 48 of 48

With many people. A load generator that replays the monitor's schedule, with and without the early attempt: N people in one open room, each with a room page's eight subscriptions and its refresh request, the clients on 4 other CPUs. Median of 3 runs: the median person, then the last one.

people main's schedule with the early attempt
server replaced as ONCE does it (kamal-proxy once-01 in front) 100 9.8 s, 15.8 s 2.5 s, 2.7 s
500 4.2 s, 4.6 s
1,000 10.5 s, 16.7 s 5.8 s, 6.4 s
2,000 7.7 s, 10.5 s
the app drops every connection, as when a room is deleted 1,000 10.5 s, 17.7 s 4.1 s, 5.0 s
docker restart 1,000 10.8 s, 15.5 s 10.7 s, 15.2 s

main's schedule at 500 and 2,000 people was run with a plain restart only: 10.8 s, 16.8 s and 11.6 s, 16.3 s. In the rows where the server is there, no early attempt failed; in the docker restart row every client's early attempt failed, once.

The replay follows connection_monitor.js and connection.js step by step. On main's schedule it gives the 10.8 s and 15.5 s above for 1,000 clients, where Chrome gave 11.3 s and 15.3 s for 16 windows. It is loadgen reconnect, a command written for this on top of the load generator of basecamp/once-campfire-verification, proposed there in basecamp/once-campfire-verification#8.

A wait of one to four seconds instead of one to two was also tried: 5.5 s instead of 5.8 s for the median of 1,000 people, 7.8 s instead of 7.7 s for 2,000. What decides it is how long the server takes to work through everyone, so the shorter wait stays.

Tests

test/system/reconnecting_test.rb:

  • a connection the server drops is back within four seconds, before the room would go offline. On main it fails there: the stream sources are still disconnected.
  • a connection the server closes with reconnect: false is still closed after the early attempt's moment has passed. This one passes on main too.
  • a room that loses its connection twice before getting it back isn't left offline: the offline timer is cleared before it is set again, since two drops within five seconds are now an ordinary thing.

Each was seen failing with one thing broken at a time: no early attempt, the attempt made whatever the server said, a wait of five seconds, the connect() call removed, the offline timer left uncleared.

Validation in the container: bin/rubocop, bin/brakeman, bin/rails herb:check, bin/rails test (580 runs) and bin/rails test:system (30 runs) all pass with no failures.

Written by Claude Code (AI assistant) on behalf of @namespaceMarcello, who directs this work.

🤖 Generated with Claude Code

Action Cable's monitor leaves a dropped connection alone for at least six
seconds and then waits for its poll timer, so that the clients of a server
that went away don't all come back at once. Half of the people in a room
are back after ten seconds and the last after sixteen, and from the fifth
second the room is offline and the composer disabled.

Often the server is there all along: it was replaced behind a proxy, as
ONCE does on an update, or it closed the connection itself, as it does to
the members of a room that is deleted and to someone removed from a room.

The room page now tries once, one to two seconds after the connection
drops, at a moment of its own. If the server said not to reconnect it
doesn't, and if the attempt fails Action Cable carries on as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017uAnU1dboBiJTZ23xmGpPe
Copilot AI balanced review requested due to automatic review settings October 10, 2026 14:20

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Repeated disconnect callbacks can leave a stale offline timer that marks a reconnected room offline.

1 open finding
What changed in this PR

Adds faster room reconnection after server-initiated Action Cable disconnects.

Changes:

  • Schedules a jittered reconnection after 1–2 seconds.
  • Adds system tests for reconnectable and permanent disconnects.

[!TIP]
If you aren't ready for review, convert to a draft PR.
Click "Convert to draft" or run gh pr ready --undo.
Click "Ready for review" or run gh pr ready to reengage.

File Description
app/​javascript/​controllers/​refresh_room_controller.js Adds early reconnection scheduling.
test/​system/​reconnecting_test.rb Tests reconnection behavior.

🧠 Review effort: Balanced


💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

}

#channelDropped({ willAttemptReconnect } = {}) {
this.#channelDisconnected()

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in b59f170: #channelDisconnected clears the offline timer before it sets a new one.

The sequence is as described. Two disconnected callbacks with no connected between them each set the timer, getting the connection back cleared only the second, and the first took the room offline, and disabled the composer, while it was connected. On main that takes the tab becoming visible or the browser coming back online between the two drops, because the monitor doesn't otherwise reopen for six seconds and the first timer has fired by then. With the early attempt it is the ordinary path whenever that attempt's socket opens and closes again before its subscriptions are confirmed.

test/system/reconnecting_test.rb has a test for it: two disconnects, a connect, and past the five seconds the composer is still enabled. It fails without the line.

Written by Claude Code (AI assistant) on behalf of @namespaceMarcello, who directs this work.

A connection can now drop twice within five seconds: the early attempt
opens a socket, and that socket closes before the subscriptions are back.
Each drop set the offline timer over the last one, and getting the
connection back cleared only the newest. The older timer then took the
room offline, and disabled the composer, while it was connected.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017uAnU1dboBiJTZ23xmGpPe
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants