Repository navigation
Reconnect soon after the server drops the cable connection - #358
namespaceMarcello wants to merge 2 commits into
Conversation
Action Cable's monitor leaves a dropped connection alone for at least six seconds and then waits for its poll timer, so that the clients of a server that went away don't all come back at once. Half of the people in a room are back after ten seconds and the last after sixteen, and from the fifth second the room is offline and the composer disabled. Often the server is there all along: it was replaced behind a proxy, as ONCE does on an update, or it closed the connection itself, as it does to the members of a room that is deleted and to someone removed from a room. The room page now tries once, one to two seconds after the connection drops, at a moment of its own. If the server said not to reconnect it doesn't, and if the attempt fails Action Cable carries on as before. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017uAnU1dboBiJTZ23xmGpPe
There was a problem hiding this comment.
🟡 Changes recommended
Repeated disconnect callbacks can leave a stale offline timer that marks a reconnected room offline.
1 open finding
What changed in this PR
Adds faster room reconnection after server-initiated Action Cable disconnects.
Changes:
- Schedules a jittered reconnection after 1–2 seconds.
- Adds system tests for reconnectable and permanent disconnects.
[!TIP]
If you aren't ready for review, convert to a draft PR.
Click "Convert to draft" or rungh pr ready --undo.
Click "Ready for review" or rungh pr readyto reengage.
| File | Description |
|---|---|
app/javascript/controllers/refresh_room_controller.js |
Adds early reconnection scheduling. |
test/system/reconnecting_test.rb |
Tests reconnection behavior. |
🧠 Review effort: Balanced
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| } | ||
|
|
||
| #channelDropped({ willAttemptReconnect } = {}) { | ||
| this.#channelDisconnected() |
There was a problem hiding this comment.
Addressed in b59f170: #channelDisconnected clears the offline timer before it sets a new one.
The sequence is as described. Two disconnected callbacks with no connected between them each set the timer, getting the connection back cleared only the second, and the first took the room offline, and disabled the composer, while it was connected. On main that takes the tab becoming visible or the browser coming back online between the two drops, because the monitor doesn't otherwise reopen for six seconds and the first timer has fired by then. With the early attempt it is the ordinary path whenever that attempt's socket opens and closes again before its subscriptions are confirmed.
test/system/reconnecting_test.rb has a test for it: two disconnects, a connect, and past the five seconds the composer is still enabled. It fails without the line.
Written by Claude Code (AI assistant) on behalf of @namespaceMarcello, who directs this work.
A connection can now drop twice within five seconds: the early attempt opens a socket, and that socket closes before the subscriptions are back. Each drop set the offline timer over the last one, and getting the connection back cleared only the newest. The older timer then took the room offline, and disabled the composer, while it was connected. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017uAnU1dboBiJTZ23xmGpPe

When the server drops a room page's cable connection, Action Cable's client doesn't try again right away. Unless the tab becomes visible or the browser comes back online in the meantime, its connection monitor never reopens a connection within six seconds of the close, and then waits for the next turn of its poll timer, so that the clients of a server that went away don't all come back at once (rails/rails#40229). A person is back after 6 to 17 seconds, half of them after about 10. From the fifth second the room is offline and the composer disabled, so everyone sees it.
The server is often there the whole time:
/up, and having kamal-proxy switch over (basecamp/once,deployWithVolume), and the proxy closes the WebSockets of the old container at once (Target#Drain: "Cancel any hijacked requests immediately"). The new server is already answering when the connections drop.reconnect: true: for every member of a room that is deleted (Room::DestroyJob) and for a person removed from a room (Membership'safter_destroy_commit).This makes the room page try once, one to two seconds after the connection drops:
refresh_room_controller.jsalready keeps the page's online and offline state, and already nudges the monitor when the browser comes back online. Itsdisconnectedcallback now also schedules oneconsumer.connect(), between one and two seconds later, at a random moment so that the people in a room don't arrive in the same instant.willAttemptReconnect: false.connect()does nothing when the connection is already open or opening, and when the attempt fails Action Cable's monitor carries on exactly as it does today.What stays as it is:
docker restart: the server is away for longer than that (4.3 s on our machine), the early attempt finds nothing, and people are back when the monitor says, as today.What it costs: the people of a room arrive within a second instead of spread over ten. A server that has just started took about 200 reconnections a second on our machine (one reconnection is the handshake, its eight subscriptions and the refresh request: 15 SELECTs and 1 UPDATE on an empty cache, 16 to 26 ms of CPU). With 1,000 people in one room
/upwent unanswered for about 5 s while they came back, with 2,000 for 5 to 8 s; under today's schedule the same work is spread out, and with 1,000 people a reconnection takes 0.06 s. In the runs below, up to 500 people in a room were all back before the composer would have been disabled.One limit of "one attempt": it is one per dropped connection. A server that accepted the early connection and dropped it again without saying not to reconnect would be tried again a second or two later, where the monitor would have waited six. A proxy with no backend refuses the upgrade instead, and a refused upgrade isn't a drop the page hears about.
Measurements
From our machine: a laptop with an AMD Ryzen 9 7940HX, Docker. The production image built from
mainat edbc779 and from this branch (Thruster, 3 Puma workers with 5 threads, 3 Resque workers, Redis inside) on 4 CPUs, the database on a Docker volume (ext4), plain HTTP, so no TLS handshakes are in these numbers.In a browser. Headless Chrome 155, 16 windows on a room, three rounds, from the socket closing to the page being online again with its refresh answered; median, lowest and highest of the 48:
mainUser#reset_remote_connections)docker restartWith many people. A load generator that replays the monitor's schedule, with and without the early attempt: N people in one open room, each with a room page's eight subscriptions and its refresh request, the clients on 4 other CPUs. Median of 3 runs: the median person, then the last one.
main's scheduleonce-01in front)docker restartmain's schedule at 500 and 2,000 people was run with a plain restart only: 10.8 s, 16.8 s and 11.6 s, 16.3 s. In the rows where the server is there, no early attempt failed; in thedocker restartrow every client's early attempt failed, once.The replay follows
connection_monitor.jsandconnection.jsstep by step. Onmain's schedule it gives the 10.8 s and 15.5 s above for 1,000 clients, where Chrome gave 11.3 s and 15.3 s for 16 windows. It isloadgen reconnect, a command written for this on top of the load generator of basecamp/once-campfire-verification, proposed there in basecamp/once-campfire-verification#8.A wait of one to four seconds instead of one to two was also tried: 5.5 s instead of 5.8 s for the median of 1,000 people, 7.8 s instead of 7.7 s for 2,000. What decides it is how long the server takes to work through everyone, so the shorter wait stays.
Tests
test/system/reconnecting_test.rb:mainit fails there: the stream sources are still disconnected.reconnect: falseis still closed after the early attempt's moment has passed. This one passes onmaintoo.Each was seen failing with one thing broken at a time: no early attempt, the attempt made whatever the server said, a wait of five seconds, the
connect()call removed, the offline timer left uncleared.Validation in the container:
bin/rubocop,bin/brakeman,bin/rails herb:check,bin/rails test(580 runs) andbin/rails test:system(30 runs) all pass with no failures.Written by Claude Code (AI assistant) on behalf of @namespaceMarcello, who directs this work.
🤖 Generated with Claude Code