Through Hyperdrive, a small fraction of queries stall for about 3.2 s or about 6.7 s when the client's extended-protocol messages for the statement (Parse + Bind + Describe + Execute + Sync) are larger than about 1 KB. The origin (PlanetScale Postgres, AWS us-east-1) executes the same statements in under 1 ms, and from the origin's side the Parse arrives late: the time is spent between the Worker's write and the origin receiving the statement.
- Statements under ~900 bytes never stall in our tests. The same query padded past ~1 KB does.
- It does not matter whether the bytes come from the SQL text or from a bound parameter.
- It is client-agnostic:
pg,postgres.jsand Effect's@effect/sql-pgall show it. - Splitting the messages into separate writes, or sending Parse + Flush and waiting for ParseComplete before sending Bind, does not help.
- The stall values are quantised (3.1–3.5 s, 6.6–7.0 s), which looks like a retransmission timeout rather than queueing.
- The rate is bursty. On the same setup we have seen 0.6–0.8% of large queries stall in one hour and 0.01–0.05% a few hours later.
We first saw this in production. CLOUDFLARE_IDS.md has the account ID, the Hyperdrive config and Worker names, and ray IDs of stalled production requests, for looking up in Cloudflare's logs. It also links the application source.
worker/src/index.js opens one pg client per request through the Hyperdrive
binding, runs 4 queries, closes the client and returns per-query timings.
?size= selects the statement size:
| size | messages | how |
|---|---|---|
small |
~520 B | SQL padded with a comment |
large |
~1470 B | same SQL, longer comment |
param |
~1375 B | shorter SQL plus one 1000-byte parameter |
All three run the same indexed lookup on a 100k-row table and return at most 5 short rows. Every statement has a unique comment, as trace comments do in real applications.
load.mjs (Node 18+, no dependencies) sends the three sizes interleaved in
random order at 64 concurrent requests, so a bad period hits all sizes equally.
-
Create the table on a Postgres origin (we used a PlanetScale Postgres development branch in AWS us-east-1):
psql "$DATABASE_URL" -f schema.sql -
Create a Hyperdrive config with caching disabled. Keep the origin connection limit below the origin's
max_connections(our branch allowed 25; the default limit exhausted it and produced unrelated connection errors):cd worker npm install npx wrangler hyperdrive create stall-repro \ --connection-string="postgres://USER:PASSWORD@HOST:5432/DATABASE" \ --caching-disabled \ --origin-connection-limit=5
To match the configuration where we first saw it, add
--ca-certificate-id=<uploaded CA id> --sslmode=verify-full. It reproduces either way (see below). -
Put the config id in
worker/wrangler.jsonc(hyperdrive[0].id). Setplacement.regionto the origin's cloud region. -
Deploy and, optionally, protect the endpoint:
npx wrangler deploy npx wrangler secret put REPRO_SECRET
-
Run the load (about 90 s):
REPRO_SECRET=... node load.mjs https://hyperdrive-size-stall-repro.<subdomain>.workers.dev
REQUESTS(default 5000 per size),CONCURRENCY(64),QUERIES(4) andSIZEScan be set in the environment. Raw per-request results go toresults.jsonl.
small has no query over 2.5 s; large and param have a few, all near 3.2 s
or 6.7 s. p50 and p99 are the same for all three sizes. Example (run 2026-09-29
20:50 UTC):
size msg bytes queries errors p50 p99 max >1s >2.5s >5s
small 519 20000 0 51 213 297 0 0 0
large 1469 20000 0 51 216 6790 10 10 3
param 1375 20000 0 49 209 3172 4 4 0
Queries over 2.5s, by size (ms):
small: none
large: 3169, 3173, 3176, 3197, 3226, 3260, 3296, 6721, 6739, 6790
param: 3145, 3153, 3155, 3172
Most stalls hit the first query on a connection, but not all.
Worker placed at aws:us-east-1, origin PlanetScale Postgres (development
branch, max_connections 25) in AWS us-east-1, 4 queries per request. Counts
are queries over 2.5 s.
| Hyperdrive config | concurrency | queries per size | small | large | param |
|---|---|---|---|---|---|
mTLS verify-full with uploaded CA, limit 5, caching off (4 runs) |
64 | 66,000 | 0 | 22 | 14 |
plain (sslmode require), limit 5, caching off (2 runs) |
64 | 40,000 | 0 | 11 | 5 |
| plain, limit 5, caching off | 128 | 20,000 | 0 | 15 | 7 |
All 74 stalls were 3126–3522 ms (58) or 6657–6955 ms (16). The slowest small
query in these runs took 461 ms. mTLS and the connection limit do not matter.
A plain config with the default origin connection limit exhausted the origin's 25 connections ("remaining connection slots are reserved", "too many clients"); that run had errors and slow queries at every size and is not this bug.
Earlier the same day (17:08–17:14 UTC), a larger harness with the same shape saw 0.6–0.8% of ~1.5 KB statements stall and 0 of 64,800 statements under ~900 bytes. Re-run at 20:45 UTC, the same harness saw no stalls in 18,000 queries, so the rate varies over time. Run the load a few times if the first run is clean.
- Origin execution time: the queries run in under 1 ms on the origin.
- The client library:
pg,postgres.jsand@effect/sql-pgbehave the same. - SQL text versus parameter: both stall once the messages pass ~1 KB.
- Write framing: one write, five writes, or Parse + Flush then waiting for ParseComplete before Bind all stall.
- mTLS / custom CA and the origin connection limit: a plain config stalls too.
- Hyperdrive caching: disabled throughout.
- Load-generator or network noise: sizes are interleaved in one run, and
smallnever stalls whilelargeandparamdo.