Repository navigation
[test][ci][tracking] Track newly observed intermittent CI failures #387
Description
Activity
SunSi12138 commented
on Aug 29, 2026 OwnerAuthorMore actionsClient-stream time-budget tests can race request registration / emission
Observed in PR #386, PR Quick #2410 at
97059f14443b167ced62db0c67eedbe5ac4a2fe2.Failures:
TimedOneWayClientStreamShouldNotStartProducerUntilRequestSurvivesEmission: expected localSharpLinkException, but the invocation completed without one.TimedClientStreamShouldNotStartProducerUntilRequestSurvivesEmission: timed out after 30s with the invocation stillWaitingForActivation.
Classification: pre-existing test synchronization race; not introduced by #386.
Evidence:
- Both failing test files are byte-identical between
dev(00e2f18c6384c785d232bd59902102d3af7ad3da) and PR feat: assembly-owned codec routing (clean port of #320) #386 head. - The previous PR head
f0bd386a8534aff77ed1086c334695f90f8bde9cpassed all 1364 unit tests under the same Ubuntu runner image / .NET SDK. - The only delta from that green head to
97059f...is a generator-only commit; it does not change request deadlines, the send queue, or send-pump timing. - In
TimedClientStreamShouldNotStartProducerUntilRequestSurvivesEmission, the test starts the async invocation and then immediately advancesManualTimeProviderand explicitly flushes. There is no deterministic barrier proving that the invocation has already registered its 5s deadline and enqueued the initial Request. If the test thread wins that race, the explicit flush can observe an empty queue; the Request is then created against the already-advanced clock and can remain behind the 10s manual-time session flush timer, matching the observed 30s wall-clock test timeout. - The one-way variant also injects clock advancement through a transport-side “next output buffer request” hook. Its failure shape is consistent with the same missing target-request synchronization around the emission boundary.
Fix direction: add deterministic synchronization tied to the target Request reaching the enqueue/emission boundary. Do not add retries or increase timeouts. I’m preparing the follow-up fix separately from #386.
SunSi12138 commented
on Aug 29, 2026 OwnerAuthorMore actionsFollow-up fix #392 has been merged into
devas3821a59f307351e6c41c7675b4ff02a2da9de005.The rerun of #386's original failed PR Quick attempt passed both client-stream time-budget tests, confirming the failures were intermittent rather than caused by #386. #392 removes the synchronization race without retries or larger timeouts by draining pre-existing connection output and advancing the manual clock at the target Request's output-buffer acquisition.
Validation before merge: formatting, Debug/Release build, and Unit Tests passed.
SunSi12138 commented
on Aug 29, 2026 OwnerAuthorMore actionsNew intermittent evidence observed while validating #386 on exact head
df5526e322f4bfb947909c521d4e4612577a94f0in PR Quick #2422:- attempt 1
quickUnit Tests: 1363/1365 passed; two existing tests hit their local ~2s timing guards:PooledAsyncStreamDispatcherLocalAbortTests.LocalAbortShouldWaitForOwnedBufferedPublicationSharpLinkClientLifecycleStateTests.StopAsyncShouldNotRunShutdownCallbacksBeforeReturning
- The test source blobs are unchanged from
dev@3821a59f307351e6c41c7675b4ff02a2da9de005(f87b573...and6ea9027...respectively). - No production/test change was made for either failure. A GitHub failed-job rerun on the same head then completed Unit Tests 1365/1365 successfully.
- The feat: assembly-owned codec routing (clean port of #320) #386 changes do not alter the
StopAsynclifecycle body; its client lifecycle diff is codec-provider construction plumbing. The LocalAbort case directly exercises the dispatcher and does not traverse the new request-stream drain API.
Current classification: scheduling-sensitive / intermittent evidence, not yet a proven #386 production regression. Per this tracker policy, do not widen timeouts or add blanket retries. If either case repeats, run focused Linux repetition/stress and separate test-barrier stabilization from production changes based on the reproduction.
- attempt 1
SunSi12138 commented
on Aug 29, 2026 OwnerAuthorMore actionsNew intermittent evidence while validating #386 on exact head
b1e0df8f477f699bdfb2b16c9a796a5d72519006, PR Quick #2463 (quickjob99133285778):Unit Tests: 1362/1364 passed; two tests hit local ~2s timeout guards.
-
SharpLinkMultiClusterClientTests.DynamicRegistrationShouldRejectASlotChangedWhileItsManifestLoadsSharpLinkMultiClusterClientTests.cs:1113- failed after ~2.01s with
TimeoutException: The operation has timed out. - I could not find an earlier issue/tracker occurrence by test name; treat this as a newly observed intermittent case, not as previously established evidence.
-
PooledAsyncStreamDispatcherLocalAbortTests.LocalAbortShouldWaitForOwnedBufferedPublicationPooledAsyncStreamDispatcherLocalAbortTests.cs:129- failed after ~2.04s with
TimeoutException: The operation has timed out. - this is a repeat of the case already recorded here from PR Quick #2422 on
df5526e..., where a same-head failed-job rerun subsequently passed Unit Tests 1365/1365 without production/test changes.
The #386 cleanup head compiled successfully in Debug and Release with 0 warnings / 0 errors before these Unit failures, and the current #386 changes are generator/routing/identity-boundary cleanup rather than the MultiCluster dynamic-registration or dispatcher timing paths.
Classification for now: scheduling-sensitive/intermittent evidence, not a proven #386 production regression. Keep these cases in #387; do not widen timeouts or add blanket retries. If
DynamicRegistrationShouldRejectASlotChangedWhileItsManifestLoadsrepeats, run focused Linux repetition/stress and determine whether it needs its own production/test synchronization issue.-
SunSi12138 commented
on Aug 29, 2026 OwnerAuthorMore actionsAdmission-control queue timing/scheduling failures during #386 validation
Observed on exact #386 head
95a03ff8b53abca7ccfcb90e0c3347abc11eb807, PR Quick #2470 (quickjob), Ubuntu 24.04 / .NET SDK 10.0.400.Unit Tests: 1359/1364 passed; five admission-control/queue cases failed:
AdmissionControlTests.ConcurrencyQueueShouldReleasePermitAndAccountingExactlyOnce— ~3.62s; assertionqueued call acquired.AdmissionControlTests.CompositeQueueRetryShouldNotConsumeAnUpstreamRatePermitTwice(fixed)— ~3.57s; assertionqueued request should reuse its previously consumed rate permit.AdmissionStateKernelTests.OverlappingGenerationsShouldShareQueueBoundsAndRetainedBytes(2, 3, queue_bytes)— ~2.51sTimeoutException.AdmissionStateKernelTests.OverlappingGenerationsShouldShareQueueBoundsAndRetainedBytes(1, 8, queue_count)— ~2.48sTimeoutException.AdmissionControlTests.QueueTimeoutShouldReleasePartitionOwnershipExactlyOnce— TUnit 30s timeout, task remainedWaitingForActivation.
I found no prior #387 entry for these names. #386 does not change the AdmissionControl/AdmissionStateKernel test files or admission-control implementation; the current work is codec generator/routing ownership cleanup. Treat this as newly observed scheduling/intermittent evidence, not as a #386 codec regression unless focused reproduction proves otherwise.
Per tracker policy: do not widen timeouts or add blanket retries. Re-run the same #386 head to unblock Generator Tests; if these recur, investigate them separately with deterministic queue/permit synchronization.
SunSi12138 commented
on Aug 30, 2026 OwnerAuthorMore actionsAdmission dynamic waiter/resize timing failures observed on PR #420
Observed while validating PR #420 on exact head
7c2cf51f8b299eca3112ac863edf4d87517fd496, PR Quick #2499 (quickjob99207503729), Ubuntu 24.04 / .NET SDK 10.0.400.Unit Tests: 1360/1363 passed; three Admission dynamic-update cases hit local timeout guards:
-
AdmissionDynamicRateLineageAndLifecycleTests.StopShouldCancelQueuedRateWaiterAndDrainRetiredTimerStateExactlyOnce- failed after ~2.17s with
TimeoutException - the test is waiting for a queued rate waiter to complete after
AdmissionStateKernel.DisposeAsync()/Stop begins draining, then verifies queue accounting, retired state, and timer disposal exactly once.
- failed after ~2.17s with
-
AdmissionDynamicRateLegacyWaiterRegressionTests.OldFixedWindowWaiterGrantShouldRemainDebtOnFastTokenBucketTarget- failed after ~2.33s with
TimeoutException - after replacing an old FixedWindow generation with a fast TokenBucket target and advancing manual time by 40s, the old captured-generation waiter did not complete within its 2s wall-clock guard.
- failed after ~2.33s with
-
AdmissionDynamicPartitionUpdateTests.PartitionConcurrencyIncreaseShouldPreserveHolderAndWakeQueuedRequest- failed after ~7.18s with
TimeoutException - the test updates partition concurrency
1 -> 2and expects the existing FIFO waiter to wake and acquire the newly available permit; the waiter did not complete in the expected window.
- failed after ~7.18s with
PR #420 changes only documentation / project-reference policy files and does not modify Admission production or test code, so these failures are not currently evidence of a #420 regression.
This is the same broad failure family as the Admission-control queue timing/scheduling cases already recorded in this tracker from PR Quick #2470, but these three exact test names were not previously listed.
Current classification: new intermittent / scheduling-sensitive Admission waiter evidence; root cause not yet established. Do not widen timeout guards or add blanket retries. If any of these repeat, run focused Linux repetition/stress with deterministic queue/timer/resize synchronization and determine whether the wakeup path exposes a production lifecycle bug or only test orchestration sensitivity.
-
SunSi12138 commented
on Aug 30, 2026 OwnerAuthorMore actionsTracker update: each remaining independent #387 root cause now has its own non-draft PR against
dev(all Ready for review):- fix(runtime): drain shared-memory final close deterministically #412 —
SharedMemoryControlChannelTests.DisposeShouldDrainTheFinalCloseSignalBeforeCompletingWakeSource - test(client): make timeout/user-cancel race deterministic #414 —
InvokeCancellableNoPayloadAsyncTimeoutAndUserCancelShouldSendSingleCancel - test(runtime): synchronize local-abort publication ownership #416 —
LocalAbortShouldWaitForOwnedBufferedPublication - test(client): start lifecycle worker before phase timeout #418 —
StopAsyncShouldNotRunShutdownCallbacksBeforeReturning - test(server): remove scheduler bound after admission permit release #421 —
OverlappingGenerationsShouldShareQueueBoundsAndRetainedBytes(both parameterized cases share this root cause) - Stabilize partition queue timeout ownership regression #423 —
QueueTimeoutShouldReleasePartitionOwnershipExactlyOnce - Stabilize concurrency queue accounting regression #424 —
ConcurrencyQueueShouldReleasePermitAndAccountingExactlyOnce - Stabilize composite admission queue retry regression #425 —
CompositeQueueRetryShouldNotConsumeAnUpstreamRatePermitTwice - Stabilize manifest-load registration race regression #426 —
DynamicRegistrationShouldRejectASlotChangedWhileItsManifestLoads
The two client-stream timing cases already fixed by merged #392 were intentionally not duplicated.
#423 needed one follow-up after CI showed the manual timer callback is allowed to complete asynchronously; the updated test now removes only the unrelated wall-clock completion bound, and its latest Unit Tests pass. #424 also received a formatting-only follow-up for the required final newline. #425 is fully green; #426 has passed formatting/build/unit/integration/native-AOT portions while the workflow finishes its remaining smoke/report tail.
Keeping #387 open until the child PRs are reviewed/merged.
- fix(runtime): drain shared-memory final close deterministically #412 —
SunSi12138 commented
on Aug 30, 2026 OwnerAuthorMore actionsAdditional unrelated intermittent failures observed while reviewing #416 / #421 / #424
While reviewing the #387 stabilization PRs, several CI failures appeared on PRs whose changes do not touch the failing path. Some related Admission cases were already tracked here, but the following exact test names were not previously recorded.
1.
RpcChannelCallShapeIntegrationTests.GeneratedProxyCallsShouldWorkWithRealRpcService(False)Observed on PR #416 head
158763c4d32ea70d9375d5630fa67b0f0b4c3c24, PR Quick #2518 (33293669721), Ubuntu 24.04 / .NET SDK 10.0.400.- Integration Tests: 409/410 passed.
- Failure: assertion
ClientStreamNoReturn totalsinRpcChannelCallShapeIntegrationTests.cs:166, reached from the generated-proxy real-service test around line 91. - test(runtime): synchronize local-abort publication ownership #416 changes only
PooledAsyncStreamDispatcherLocalAbortTeststest synchronization; it does not modify RPC generated-proxy call-shape production or integration-test code.
Classification: new intermittent integration evidence; unrelated to #416; root cause not yet established. If it repeats, investigate client-stream no-return completion/accounting synchronization rather than widening the eventual-consistency timeout.
2.
AdmissionDynamicPartitionRateTransitionTests.LateOldPartitionFixedWindowGrantShouldRemainDebtOnTokenBucketTargetObserved on PR #421 head
5d824a16e10a1009dc2b05bf9ac59dca4041af33, PR Quick #2497 (33292604360).- Failed after ~2.35s with local
TimeoutExceptionaroundAdmissionDynamicPartitionRateTransitionTests.cs:175. - test(server): remove scheduler bound after admission permit release #421 is test-only and changes only
AdmissionStateKernelTests.OverlappingGenerationsShouldShareQueueBoundsAndRetainedBytes; it does not change this test or Admission production code. - This is related to the broader Admission generation/waiter timing family already seen in this tracker, but this exact test name had not been listed.
Classification: new scheduling-sensitive Admission generation/waiter evidence; unrelated to #421.
3.
AdmissionDynamicPartitionUpdateTests.SelectorReplacementShouldKeepOldQueuedRequestOnOldNamespaceObserved in the same PR #421 / PR Quick #2497 run.
- Failed after ~7.21s with
TimeoutExceptionaroundAdmissionDynamicPartitionUpdateTests.cs:180. - test(server): remove scheduler bound after admission permit release #421 does not touch this test or the production selector-replacement path.
Classification: new scheduling-sensitive Admission selector/queued-waiter evidence; unrelated to #421; root cause not yet established.
Note: the same run also failed
ConcurrencyQueueShouldReleasePermitAndAccountingExactlyOnce, but that exact case was already recorded in #387 and was subsequently addressed by #424, so it is not duplicated here.4.
UnsizedStreamingPreCreditTests.CreditStarvationShouldBoundLongLivedSerializedOwnersAndWaitersObserved on PR #424 head
f1e1ca08111d311633ca3c7d86beb1eeef157659, PR Quick #2514 (33293525861).- Failed after ~2.11s with local
TimeoutExceptioninExpectSameException, aroundUnsizedStreamingPreCreditTests.cs:153/ test line 84. - Stabilize concurrency queue accounting regression #424 is test-only and changes one Admission concurrency-queue test to use
ManualTimeProvider; it does not touch streaming pre-credit production or tests.
Classification: new scheduling-sensitive streaming pre-credit evidence; unrelated to #424; root cause not yet established.
Note: the same run also failed the three
CompositeQueueRetryShouldNotConsumeAnUpstreamRatePermitTwicevariants; that exact family was already recorded here and addressed by #425, so it is not duplicated.Per #387 policy: keep these as evidence, do not add blanket retries or simply enlarge wall-clock guards. If they recur, use focused Linux repetition/stress and deterministic phase barriers to separate test orchestration sensitivity from a production concurrency/lifecycle defect.
SunSi12138 commented
on Aug 30, 2026 OwnerAuthorMore actionsFollow-up on the three Admission dynamic waiter/resize failures from PR Quick #2499:
- test(server): remove scheduler bound after legacy rate grant #427 —
AdmissionDynamicRateLegacyWaiterRegressionTests.OldFixedWindowWaiterGrantShouldRemainDebtOnFastTokenBucketTarget- root cause classified as a test-only scheduler guard after
ManualTimeProvider.Advance(40s)has already triggered the source-generation grant - removes only the unrelated 2s wall-clock
WaitAsync; transition-debt/expiry assertions are unchanged
- root cause classified as a test-only scheduler guard after
- test(server): remove scheduler bounds from admission stop drain #428 —
AdmissionDynamicRateLineageAndLifecycleTests.StopShouldCancelQueuedRateWaiterAndDrainRetiredTimerStateExactlyOnce- root cause classified as test-only ThreadPool continuation latency after synchronous kernel draining cancellation
- removes the two local 2s scheduler guards while preserving shutdown error, queue accounting, state reclamation, and timer-disposal assertions
- test(server): remove scheduler bound after partition resize #429 —
AdmissionDynamicPartitionUpdateTests.PartitionConcurrencyIncreaseShouldPreserveHolderAndWakeQueuedRequest- root cause classified as test-only scheduler latency after the synchronous 1 -> 2 partition concurrency target commit has woken the queued waiter
- removes only the post-resize 2s wall-clock guard; namespace/holder/FIFO/accounting assertions are unchanged
Each case is isolated in its own non-draft PR against
dev; CI is running.#412 was also revised after review found a race in its first fix: the separate
_writerActive == 0/ pending-empty observation could be invalidated by a concurrent outbound signal before Close enqueue, allowing an earlier blocked write to bypass the 250ms cleanup path. The updated #412 serializes writer active state, pending enqueue/dequeue, and Dispose'sidle + empty -> Closetransition under one outbound-state gate. Unbounded waiting for Close-start is now used only when Dispose atomically owns a genuinely idle/empty writer; otherwise the original bounded cleanup path is preserved.- test(server): remove scheduler bound after legacy rate grant #427 —
SunSi12138 commented
on Aug 30, 2026 OwnerAuthorMore actionsAdditional unrelated intermittent failure observed while reviewing #433
SharpLinkClientRetryTests.HugeBuiltInJitteredRetryDelayShouldRemainCancellableObserved on PR #433 head
b8f493a927b0e935831707a51e052f656953d5c7, PR Quick #2553 (33295641662), Ubuntu 24.04 / .NET SDK 10.0.400.- Unit Tests: 1362/1363 passed.
- Failure: local
TimeoutExceptionafter ~4.22s atSharpLinkClientRetryTests.cs:173-174. - test(server): remove scheduler bound after old partition rate grant #433 changes only
AdmissionDynamicPartitionRateTransitionTests.LateOldPartitionFixedWindowGrantShouldRemainDebtOnTokenBucketTarget; it does not touch client retry/jitter/cancellation production or test code. - Formatting, maintainability gates, Debug/Release builds, generated-assembly dependency checks, CodeQL, and codec-compatibility jobs passed before/alongside the unrelated unit failure.
Classification: new scheduling/timing-sensitive client retry cancellation evidence; unrelated to #433; root cause not yet established.
Per #387 policy, keep this as evidence rather than widening the local timeout or adding blanket retries. If it repeats, use focused repetition with a deterministic barrier around cancellation registration / retry-delay scheduling.
SunSi12138 commented
on Aug 30, 2026 OwnerAuthorMore actionsTracker follow-up for the two currently unchecked cases in #387:
-
test(integration): remove wall-clock bound from oneway completion poll #436 —
RpcChannelCallShapeIntegrationTests.GeneratedProxyCallsShouldWorkWithRealRpcService(False)/ClientStreamNoReturn totals- root cause: the exact total includes
[Oneway]client-stream handlers, whose server-side completion is not implied by client-side return; the helper imposed an unrelated 1s system-clock deadline on eventual server completion - fix keeps the exact total assertion and polling, removes only that local wall-clock deadline
- non-draft / Ready for review; PR Quick and CodeQL are green
- root cause: the exact total includes
-
test(client): synchronize huge retry backoff before stop #439 —
SharpLinkClientRetryTests.HugeBuiltInJitteredRetryDelayShouldRemainCancellable- root cause:
Task.Delay(20)guessed that the built-in huge jittered retry backoff had already been scheduled beforeStopAsync - fix uses
ManualTimeProvidertimer registration as a deterministic phase barrier before Stop, while retaining all 32 iterations, the original 2s cancellation/Stop bounds, and the expectedConnectionClosedresult - non-draft / Ready for review; PR Quick and CodeQL are green
- root cause:
Both fixes are test-only; neither changes production RPC/retry/timeout semantics. Keeping #387 open for review/merge and future intermittent evidence.
-
SunSi12138 commented
on Aug 30, 2026 OwnerAuthorMore actionsSendPumpProgressIsolationTests.ProgressInterleaveServesProgressWhileNormalQueueNeverEmpties现象:
- PR ci: implement bounded PR Fast gate #437 最新 head
dc84cc42e1c949e8e122a8211a2271461a9100a2的 PR Quick run #2573(run33297143485)在Unit Tests中偶发失败。 - 失败信息:
the progress frames must drain as one contiguous batch (indices 0,1,2,3,4,5,6,7,72,73)。 - 同一 PR merge commit / 同一组 Unit Tests 在并行的
PR Fast / fastrun perf: recover v0.7.2 unary throughput and allocations #3(run33297143414)中通过,说明当前证据更符合 scheduling-sensitive / intermittent failure,而不是 ci: implement bounded PR Fast gate #437 引入的确定性回归。
证据 / run:
- PR Quick: https://github.com/SunSi12138/SharpLink/actions/runs/33297143485/job/99218706198
- PR Fast: https://github.com/SunSi12138/SharpLink/actions/runs/33297143414/job/99218684594
- PR: ci: implement bounded PR Fast gate #437
环境:
- Ubuntu 24.04 hosted runner
- .NET SDK 10.0.400
- PR Quick 失败时 Unit Tests: 1362 passed / 1 failed
初步关注点:
- 该测试依赖“progress frames 必须作为一个连续 batch 出现在 wire order 中”的调度假设;本次实际观察到 8 个 Ping 在开头、剩余 2 个 Ping 在 index 72/73,说明 pump/pipe 调度可以把这 10 个 progress frame 分成多个合法可见批次。
- 历史 test(runtime): stabilize progress starvation regression #291 曾修过同一
SendPumpProgressIsolationTests文件中的另一个 scheduler-dependent starvation regression test,但不是这个具体 case;本条按新 flake 单独记录。 - 按 tracker 原则,优先做 focused Linux repetition/stress,并确认该 contiguous-batch assertion 是否真的是 production contract;不要通过 blanket retry、扩大 timeout 或简单弱化 correctness assertion 处理。
- PR ci: implement bounded PR Fast gate #437 最新 head
SunSi12138 commented
on Aug 30, 2026 OwnerAuthorMore actionsNew intermittent evidence from #439 / PR Quick #2613
Observed on #439 head
0d86048f60aa20fcb56e138c88bf6ba4ebede139, PR Quick #2613 (33298940958), Ubuntu 24.04.4 / .NET SDK 10.0.400.Failure:
TransportCleanupTests.StreamConnectionDisposeShouldWaitForOutstandingReadRelease- Unit Tests: 1362/1363 passed; this was the only failure.
- ~2.781s
TimeoutExceptionatTransportCleanupTests.cs:44.
The timeout is the second local 2s guard: after disposal has requested reader completion, the test verifies disposal is still pending while the consumer owns the
ReadResult, callsreader.AdvanceTo(result.Buffer.End)to release that ownership, and then times out waiting forconnection.DisposeAsync()to finish.Classification: new intermittent / lifecycle-sensitive evidence; unrelated to #439, root cause not yet proven test-only.
Evidence for unrelatedness:
- test(client): synchronize huge retry backoff before stop #439 changes only
SharpLinkClientRetryTests.HugeBuiltInJitteredRetryDelayShouldRemainCancellable. TransportCleanupTests.csis byte-identical between test(client): synchronize huge retry backoff before stop #439 base0edd85c7...and head (f6fc72e0530d9e4af673971df15771f22dafbae5).TransportConnection.cs, includingStreamTransportConnectionandReadOwnershipPipeReader, is also byte-identical between base and head (aa7de7fbbaf17cea3e6cb25d90176237117fb7f5).- test(client): synchronize huge retry backoff before stop #439's target retry test did not fail in #2613.
The production path has deterministic read-release signaling, but completion after release still crosses asynchronous continuations (
RunContinuationsAsynchronously/Task.Yield) and the underlyingPipeReader.CompleteAsync. One CI occurrence is not enough to decide whether the local 2s guard is merely scheduler-sensitive or whether there is a latent lifecycle/completion race.Next step should be focused Linux repetition/stress of this exact test and inspection of the post-
AdvanceTocompletion path. Do not widen the timeout or add blanket retries as the fix.SunSi12138 commented
on Aug 30, 2026 OwnerAuthorMore actionsSharpLinkServerHostedServiceTests.UnexpectedSuccessfulRunCompletionShouldStopTheHostObserved in PR #445,
PR Fastrun33299359487, merge commit477142fd1e072a41012d7606b2036a00599f5edb(headae8c615239f830cc11028b395ecda967373756da).现象:
- Unit Tests: 1362/1363 passed; only this test failed.
- Failure assertion:
an unexpected successful Server run-loop exit must stop the owning Host. - Test elapsed ~608ms and uses
Task.WhenAny(lifetime.StopRequested.Task, Task.Delay(500ms))as a wall-clock guard after callingserver.StopAsync(TimeSpan.Zero).
证据 / run:
- ci: stop PR Quick from running on every pull request #445 changes exactly one workflow line: removes
pull_requestfrom.github/workflows/pr-quick.yml; it does not modify Hosting, Server, or test code. - Restore, formatting, maintainability gates, Release build, and generated-assembly dependency verification all passed before Unit Tests.
SharpLinkServerHostedService.ObserveRunTaskAsyncobserves the server run task asynchronously and callsapplicationLifetime.StopApplication()when it sees an unexpected successful completion. The test has no deterministic barrier for that continuation; it only waits up to 500ms wall-clock.- Runner: Ubuntu 24.04, .NET SDK 10.0.400.
初步关注点:
- Strong candidate for test synchronization / scheduler-latency flake rather than a ci: stop PR Quick from running on every pull request #445 regression.
- The 500ms wall-clock
Task.Delayguard can lose under hosted-runner / parallel-test scheduling pressure even when the production continuation is correct. - Do not fix by merely widening the timeout or adding blanket retry. Prefer a deterministic observation/barrier around
StopApplication()/ run-task completion if reproduction confirms this classification.
状态: 待复现 / 待归因。A rerun of the exact failed Fast job has been started; record the rerun result here once complete.
关联 PR: #445
SunSi12138 commented
on Aug 30, 2026 OwnerAuthorMore actionsNew PR Fast unit-test flakes observed on #445
Observed in PR #445, PR Fast run
33299548254on merge commit460c9369994bdb0d82400d9b7a937f940b44f2e3(head6caa29679b63f1cc78d123b68b80edea424b314f, base8238f93d86016af2258da9ef9ce7429a9888e0e1), Ubuntu 24.04 / .NET SDK 10.0.400.Unit Tests result: 1361/1363 passed; two failures:
-
SendPumpProgressIsolationTests.ProgressInterleaveServesProgressWhileNormalQueueNeverEmpties- Failed in 43ms.
- Assertion expected all 10 progress Ping frames to form one contiguous batch.
- Observed Ping indices:
0,1,2,3,68,69,70,71,72,73. - The test enqueues 130 normal frames and then sequentially enqueues 10 Ping frames with no deterministic barrier preventing the send pump from waking after only part of the Ping producer loop has completed.
- Production send-pump policy drains progress at the loop top and every 64 normal frames. The observed split is consistent with a legal schedule: the pump drains the first 4 Pings, processes 64 normal frames, then drains the remaining 6 at the interleave boundary.
- The test file is byte-identical on ci: stop PR Quick from running on every pull request #445 base/head (
6a0cdfd79df65768e8ac9a141d5f81b5c954e865). ci: stop PR Quick from running on every pull request #445 itself only changes thePR Quickworkflow trigger. - Initial classification: test synchronization / scheduling assumption, not evidence of a ci: stop PR Quick from running on every pull request #445 production regression. The intended invariant should be clarified: bounded progress service is guaranteed, but atomic contiguity of concurrently-enqueued progress frames is not obviously a runtime contract.
-
NegotiatedSessionOptionsTests.DrainingShouldRejectNewRequestsAndPreserveExistingCallFrames- Failed with
TimeoutExceptionafter ~3.27s atFlushSendQueueAsync().AsTask().WaitAsync(TimeSpan.FromSeconds(2)). - Test source is byte-identical on ci: stop PR Quick from running on every pull request #445 base/head (
4b5c0e2cf4721658f82bb3626ab66f9a527028c6). - Initial classification: scheduler/lifecycle-sensitive evidence; root cause not yet established. Focused repetition should determine whether this is send-pump flush continuation latency or a real draining/flush race.
- Failed with
Both failures are unrelated to #445's one-line workflow-trigger diff. Per #387 policy, do not fix either by merely increasing timeouts, blanket retry, or weakening the correctness assertion. Prefer deterministic synchronization / focused stress and confirm the actual production invariant first.
-
79 remaining items
- added a commit that references this issue
on Sep 18, 2026 SunSi12138 commented
on Sep 18, 2026 OwnerAuthorMore actions2026-09-18 update: Nightly Regression #284/#285 follow-up has been resolved and merged into
dev.- fix: close deadline arm race and stabilize #387 lifecycle regressions #698 (
e355e0f3101756f94fbae8f0ee45ad0bb22061d9, squash) fixes the production absolute-deadline timer arm TOCTOU and adds deterministic full/partial arm-drift plus caller-cancellation arbitration regressions. It also replaces the disconnected-cleanup test's premature ownership assertion with the framework-supervisor completion barrier, and removes the per-round 2s remote stream-disposal wall-clock guard while retaining the enclosing test timeout. - test(client): synchronize dynamic reconnect replacement timer #699 (
e1edc26b292ee0c464f0f94aa31895c45e93b21c, squash) fixes the macOS dynamic reconnect replacement flake by waiting for the exact live 1s reconnect timer before advancing manual time; production reconnect semantics are unchanged.
Both PRs were reviewed with no remaining correctness blockers and their required
fastchecks passed before merge. The tracker remains open for future intermittent CI failures.- fix: close deadline arm race and stabilize #387 lifecycle regressions #698 (
SunSi12138 commented
on Sep 18, 2026 OwnerAuthorMore actions2026-09-18 follow-up: #701 has been merged into
devasa2d1ee3b6e330646fc2677bdf489cddcfdcd6698.The Browser codec evidence policy is now narrowed to the documented compatibility boundary: only Browser <-> hosted-desktop failures in the existing
auto-layout-release-scopedcategory are retained as non-blocking evidence when they have one of the recognized layout/decode mismatch shapes. This covers the known Mono/wasm32 vs CoreCLR/64-bit managed-layout variance that was blocking the main Release Gate.Browser self-roundtrip, desktop <-> desktop compatibility, every other fixture category, fixture/identity/hash/completeness checks, and unexpected/malformed classifications remain blocking. No production codec hot-path change, retry, timeout widening, or blanket continue-on-error was introduced.
#700 has been refreshed to the new
devtree and will rerun the Release Gate.SunSi12138 commented
on Sep 18, 2026 OwnerAuthorMore actions2026-09-18 Browser evidence follow-up: #702 merged into
devas62d876d01321c0afac22b56c54a5b40039180481.The first #701 Release Gate rerun confirmed the intended C# policy behavior: the known Browser -> desktop auto-layout rows were emitted with
blocking=falseand the verifier reportedblocking failures: 0. The remaining failure was only the JS strict evidence-shape checker: forDESERIALIZE_REJECTED,Fixture<T>.Verifycan retain a previously computedlogicalEquality=falsewhen diagnostic value rendering throws.#702 narrows that validator mismatch by allowing exactly
nullorfalsein that field for the already evidence-only rejected row;true, other classifications, other fixture categories, and other evidence invariants remain rejected. #700 has been refreshed again to the currentdevtree.SunSi12138 commented
on Sep 18, 2026 OwnerAuthorMore actions2026-09-18 Release Gate #606 Windows integration flake:
- Run
35336209933, job105571522074 SharedMemoryListenerShouldRejectBadHandshakesAndAcceptNextClientfailed after the three bad-handshake cases withShared-memory transport handshake timed out after 00:00:00.1000000on the final healthy client.- The listener's 100ms handshake timeout starts only after a pipe connection is accepted; the client factory's same 100ms timeout starts before
ConnectAsync, so it also charges the scheduler/accept-loop tail while the listener cycles from the prior rejected connection. - The server-side 100ms rejection contract remains valid; the flake is caused by reusing that narrow server test budget for the final healthy recovery client's local end-to-end wait.
A focused test-only follow-up is being prepared to keep the listener at 100ms while giving only the final healthy client its own recovery timeout. No production transport behavior or bad-handshake timeout is being widened.
- Run
SunSi12138 commented
on Sep 18, 2026 OwnerAuthorMore actions2026-09-18 follow-up: #703 and #704 are now merged into
dev.- test(runtime): separate shared-memory recovery client timeout #703 (
c19bbaac69585c2b60bf42009a345e8c2a50760b) separates the final healthy shared-memory recovery client's local timeout from the server-side 100ms bad-handshake contract. - test(codec): add Browser CoreCLR compatibility evidence #704 (
6ba3a8a4e6ab55cc8c0548de689b04aae0b3111a) adds experimental Browser/WASM CoreCLR codec evidence using .NET 11 RC1 withUseMonoRuntime=false. The new graph checks six desktop producers + Browser Mono + CoreCLR self in Browser CoreCLR and checks the CoreCLR Browser producer from all six desktop consumers.
The Browser CoreCLR graph is intentionally non-blocking while the runtime is pre-GA. Existing Mono Browser evidence and the release-gated desktop contract remain unchanged. #700 has been refreshed to the new
devtree so Release Gate can execute the new evidence lane.- test(runtime): separate shared-memory recovery client timeout #703 (
SunSi12138 commented
on Sep 18, 2026 OwnerAuthorMore actions2026-09-18 Browser CoreCLR identity follow-up: #705 merged into
devas400f490a4ca18223ee6acf3b5f6c742ee2e4987b.Release Gate #607 established that .NET 11 RC1 Browser CoreCLR successfully publishes and runs in headless Chrome with
UseMonoRuntime=false. It also exposed thatType.GetType("Mono.Runtime")returns null on the existing Browser/Mono probe, so reflection cannot be used to distinguish the Browser runtime family.#705 now records Browser runtime family from the actual project runtime selection: default net10 Browser remains
Mono/platform-runtime-pack, while the explicitUseMonoRuntime=falsebuild recordsCoreCLR/build-runtime-selection. Observed TFM/RID/pointer size/framework/runtime/SDK/compilation mode remain runtime evidence.#700 has been refreshed again so the full evidence graph can be rerun with corrected identities.
SunSi12138 commented
on Sep 18, 2026 OwnerAuthorMore actions2026-09-18 final follow-up: the CI stability / Browser codec work has now landed on
main.- Release Gate docs: make README NuGet-first with runnable templates #608 completed with
release-summarysuccess. - Existing Browser Mono evidence passed.
- Browser CoreCLR forward evidence passed with the corrected deterministic runtime identity.
- Browser CoreCLR -> all six desktop reverse verification lanes passed.
- Desktop compatibility summary passed.
publish-nugetwas skipped.- chore: sync reviewed CI stability fixes into main #700 merged into
mainase46d26b20e2036832f345f7b354dc6575a56a335.
No new tag or GitHub Release was created; the latest published release remains
v2.0.0. The tracker remains open for future intermittent CI failures.- Release Gate docs: make README NuGet-first with runnable templates #608 completed with
SunSi12138 commented
on Sep 19, 2026 OwnerAuthorMore actions2026-09-19 documentation follow-up: codec compatibility terminology is now aligned with the current CI evidence.
- CoreCLR is the primary compatibility matrix.
- Mono/AutoLayout-only mismatches are not part of the CoreCLR release contract and should not be treated as release blockers unless the same failure appears on a contracted CoreCLR edge.
- docs(codec): sync CoreCLR-first compatibility contract #708 passed Release Gate [Runtime Configuration][Call Capacity] Runtime resize for per-connection and server-wide active RPC limits #611 and merged to
mainas8cbcffaf466f13ef546763b73ba68d37e545b579.
This is documentation/tracking alignment only; no production or test-gating behavior was changed.
SunSi12138 commented
on Sep 19, 2026 OwnerAuthorMore actions新增 intermittent CI failure —
StaticClusterReconnectShouldBeSingleFlightAtTheProviderBoundary现象:
- PR feat: add optional generation control extension #706 在最终 head
c8f8b21219643b5614055a1adaa609eeea960485上,一轮由 PR description edit 触发的 PR Fast 中 Unit Tests 失败。 - 1851 个 Unit Tests 中仅此 1 个失败。
- 失败信息:
endpoint static-first has no active reconnect owner。 - 栈位于
SharpLinkClientLifecycleReconnectSupport.GetStaticReconnectTask(...):71,由SharpLinkClientLifecycleReconnectTests.StaticClusterReconnectShouldBeSingleFlightAtTheProviderBoundary():150触发。
证据 / run:
- 失败 run: PR Fast
35441391215,首次fastjob105892663366。 - 同一 head 的前一轮 PR Fast
35441150436已完整通过。 - 对失败 job 原样重跑后,新
fastjob105893110493成功,Unit Tests、Generator Tests 及后续 Fast checks 全部通过。 - 同一 head 的 PR Extended
35441347233也成功,包括 Integration Tests、NativeAOT smoke、package smoke、demo/load 及桌面 codec compatibility。
环境 / 变更相关性:
- 该测试文件
test/SharpLink.UnitTests/Client/SharpLinkClientLifecycleReconnectTests.cs与 helperSharpLinkClientLifecycleReconnectSupport.cs在 PR head 与当前dev上 blob SHA 完全一致。 - PR feat: add optional generation control extension #706 不修改 Client reconnect production path 或上述测试。
- 因此当前证据更符合 existing timing / reconnect-owner observation flake,而不是 GenerationControl 回归。
初步关注点:
- 测试在读取
StaticClusterRuntimereconnect owner 时,可能存在“owner 已被消费/完成但测试尚未观察到下一 owner”或 provider-boundary single-flight 状态切换的短暂窗口。 - 若后续复现,优先给 reconnect owner publication / timer arm / ownership handoff 增加确定性 barrier,再做 focused repetition;不要用 blanket retry、sleep 或单纯扩大 timeout 作为最终修复。
状态: 待复现 / 待归因。
关联:
- PR feat: add optional generation control extension #706
- final head:
c8f8b21219643b5614055a1adaa609eeea960485
- PR feat: add optional generation control extension #706 在最终 head
SunSi12138 commented
on Sep 19, 2026 OwnerAuthorMore actions新增 intermittent CI failure —
TimedBatchShouldExtendBatchForFrameArrivingBeforeDeadline现象:
- Release Gate run
35448326880的matrix-build-test (macos-latest)唯一 Unit Test 失败。 - 1851 个 Unit Tests 中仅此 1 个失败。
- 失败信息:
TimeoutException: The operation has timed out. - 栈定位到
test/SharpLink.UnitTests/Runtime/SendPumpTests.cs:174/:188。 - 失败用时约 2.18s,属于测试内 2s timing guard 超时。
证据 / run:
- Release Gate run:
35448326880 - failed job:
105911035576 - 环境:macOS arm64 Release Unit Tests
- 同一 merge commit 的 PR Fast、PR Package Smoke、CodeQL、Codec Padding Security Evidence、Nightly Regression 均成功;Release Gate 其余 Windows/Linux matrix、AOT、chaos、pack、codec lanes 也通过。
变更相关性:
- PR feat: add optional generation control extension #706/feat: add generation control extension #717 不修改
SendPumpTests.cs。 - 该文件在当前
dev与 generation-control 分支 blob SHA 均为f69d368a08a9d699268ec7d68ee1775e5cf79ab6,完全一致。 - 因此当前证据不支持 GenerationControl 回归,先按独立 timing-sensitive flake 跟踪。
初步关注点:
- 检查测试对 timed-batch deadline / second-frame arrival 的同步是否依赖 wall-clock scheduler timing。
- 若可复现,优先引入 deterministic synchronization/manual-time seam,而不是扩大 timeout 或 blanket retry。
状态: 待复现 / 待归因。
关联:
- PR feat: add optional generation control extension #706 / replacement PR feat: add generation control extension #717
- merge commit
dfc123fb00d5dae162000068385180326f5c7e6d
- Release Gate run
SunSi12138 commented
on Sep 19, 2026 OwnerAuthorMore actionsFollow-up for
TimedBatchShouldExtendBatchForFrameArrivingBeforeDeadline:- the exact failed Release Gate macOS job was re-run unchanged
- re-run job:
105914248807 - result: success
- Unit Tests passed, including the previously failing timed-batch case
- Generator Tests, Load Test Tests, Integration Tests, Pack, and Package Smoke also passed
No code or timeout threshold was changed. This confirms the observed failure was intermittent on the same code/commit; keep the case tracked for deterministic synchronization/root-cause work rather than treating it as a GenerationControl regression.
SunSi12138 commented
on Sep 21, 2026 OwnerAuthorMore actions新增 intermittent CI failure — allocation gate stability spread
现象:
- PR perf: hoist streaming codec capability checks to stream scope #724 final head
b4a311662dbc2250d42a2e3331aa65dbdf898703的 PR Fast run35616693360中,fastjob 成功,只有allocation-gatejob106389543642失败。 - 失败 case:
rpc-add-sharedmemory-c1。 - observed median:
1324.768 B/op,低于 allowed median1450.000 B/op。 - observed spread:
53.558 B/op,略高于 allowed spread50.000 B/op,因此 stability check 失败。 - 同一 gate 的
rpc-add-sharedmemory-c8、rpc-oneway-sharedmemory-c1、send-pump-idle-wake-balanced均通过;allocation gate harness self-tests 也通过。 - 同一 head 的 PR Package Smoke 和 CodeQL 成功。
变更相关性:
- PR perf: hoist streaming codec capability checks to stream scope #724 修改的是 streaming codec capability hoisting 与对应 regression tests,不涉及
rpc-add-sharedmemory-c1的 unary/shared-memory allocation path。 - 当前失败不是 median allocation budget regression,而是单次 sample spread 超阈值 3.558 B/op,初步按 runner/perf-noise stability flake 跟踪。
证据 / run:
- PR Fast:
35616693360 - failed job:
106389543642 - PR: perf: hoist streaming codec capability checks to stream scope #724
- head:
b4a311662dbc2250d42a2e3331aa65dbdf898703
状态: 已记录;原样重跑 failed job 验证。
- PR perf: hoist streaming codec capability checks to stream scope #724 final head
- added a commit that references this issue
on Oct 9, 2026 SunSi12138 commented
on Oct 9, 2026 OwnerAuthorMore actionsAllocation c1/c8:记录调查结论,维持未解决
本轮按维护者决定结束调查,暂不修复,也不继续推进门禁拆分或移除 SharedMemory 的方案。生产实现、预算和正式门禁判定保持原样;
rpc-add-sharedmemory-c1/c8仍为未解决,本 issue 保持开放。已验证的机制
- 范围是 Linux x64、.NET 10.0.12;诊断基线为
e91f82118bb4d67997bc76b8a579d51460a28c1b。本次观测发生在 SharedMemory 的通知管道。 - 预热后,底层管道读取同步完成时,调用边界分配为 0 B;确实挂起时为 144 B。保留的 trace 采到了
PipeStream.ReadAsyncCore对应的AsyncStateMachineBox<int, ...>,采样对象大小均为 144 B;独立强制路径实验验证了相应总分配。这是保存异步执行状态的 Task-backed 对象,不是读取的 1 字节缓冲区。默认async ValueTask不自动池化。 - 在一次完整的 c1 观测中,低/高样本差为 21.504 B/op;多出 600 次未完成的 control-pipe read,少一次 SharedMemory 挂起读取,对应调用边界差 21.520 B/op,残差变化 −0.016 B/op。总 control read 只增加 134 次,说明变化主要体现为同步/挂起的完成比例,不能把 600 全部解释成额外通知。
- 最后的固定混合实验中,Memory/ValueTask 与 byte-array/Task 两个公开重载在 0/25/50/75/100% 挂起比例下,均测得 0/36/72/108/144 B/read;取消、取消后复用及 pending-read disposal 检查均通过。换公开重载没有消除这项开销。
- perf(runtime): make the suspended-read wrapper allocation-free (#740) #753 处理的是 SharpLink 外层
ReadOwnershipPipeReader的约 248 B 包装分配;这里的 144 B 位于 .NET 内部,是不同层次。perf(runtime): make the suspended-read wrapper allocation-free (#740) #753 不会直接消除它。这些数值也不能外推到 Windows 或所有 transport。
证据与边界
- 校准后的 trace / 原始门禁对照
- 固定路径计数与 SharedMemory ready/pending 对照(完整证据)
- 最后一次公开 PipeStream API 校准(完整证据)
- .NET 10.0.12 PipeStream 源码;默认 ValueTask builder。
上述固定诊断序列保留全部样本,来源和结果已独立复核;+512 B/operation 的负向控制仍正确拒绝 c1/c8。诊断中的真实 RPC 门禁没有重现历史 spread 超限,最后一次校准也没有运行 RPC 门禁。路径计数只覆盖同步调用阶段,不能当作完整异步成本模型;历史超限样本没有对应计数,因此尚未完整证明历史失败根因,不能据此判定那些失败无害或已修复。本轮未找到安全、证据充分的最小修复,并不等于偶发分配原则上无法检测。
另行完成的
TimedBatchShouldExtendBatchForFrameArrivingBeforeDeadline已由 #794 合入dev(72806161a0580a047f79348b72783025501daf79);这是独立的 test-only 修复,不解决这里的 allocation c1/c8。- 范围是 Linux x64、.NET 10.0.12;诊断基线为
背景
用于持续记录和归因当前
dev/ PR CI 中新出现的偶发测试失败,方便后续集中复现、定位并修复。这不是 blanket retry / quarantine 列表。每个 case 最终都应尽量判断为:
历史上一批 timing / runner-sensitive failures 见 #275;本 issue 从当前新观察到的 case 开始继续跟踪,后续可以直接追加新条目。
当前状态
已处理 / 已合并
TimedBatchShouldExtendBatchForFrameArrivingBeforeDeadline— test-only 同步修复已由 test: synchronize timed send-batch deadline assertions #794 合入dev(72806161a0580a047f79348b72783025501daf79);与下述 allocation c1/c8 两项独立。TimedOneWayClientStreamShouldNotStartProducerUntilRequestSurvivesEmission— 已由 test: make client-stream time-budget emission tests deterministic #392 修复并合入dev。TimedClientStreamShouldNotStartProducerUntilRequestSurvivesEmission— 已由 test: make client-stream time-budget emission tests deterministic #392 修复并合入dev。SharedMemoryControlChannelTests.DisposeShouldDrainTheFinalCloseSignalBeforeCompletingWakeSource— 已由 fix(runtime): drain shared-memory final close deterministically #412 修复并合入dev;最终版本将 writer active state、pending enqueue/dequeue 与 Dispose 的idle + empty -> Close决策串行化,修复首版 review 发现的 shutdown race。InvokeCancellableNoPayloadAsyncTimeoutAndUserCancelShouldSendSingleCancel— 已由 test(client): make timeout/user-cancel race deterministic #414 修复并合入dev。PooledAsyncStreamDispatcherLocalAbortTests.LocalAbortShouldWaitForOwnedBufferedPublication— 已由 test(runtime): synchronize local-abort publication ownership #416 修复并合入dev。SharpLinkClientLifecycleStateTests.StopAsyncShouldNotRunShutdownCallbacksBeforeReturning— test(client): start lifecycle worker before phase timeout #418 先将 dedicated worker 启动改为确定性;test(client): remove scheduler bound from shutdown callback dispatch #435 进一步移除 cancellation callback dispatch 后残留的本地 scheduler guard,均已合入dev(test(client): remove scheduler bound from shutdown callback dispatch #435 squash3e772ef0de75bda9f79d5f3508ef01f573bb6254)。AdmissionStateKernelTests.OverlappingGenerationsShouldShareQueueBoundsAndRetainedBytes— 两个参数化 case 共用同一 root cause,已由 test(server): remove scheduler bound after admission permit release #421 修复并合入dev。AdmissionControlTests.QueueTimeoutShouldReleasePartitionOwnershipExactlyOnce— 已由 Stabilize partition queue timeout ownership regression #423 修复并合入dev。AdmissionControlTests.ConcurrencyQueueShouldReleasePermitAndAccountingExactlyOnce— 已由 Stabilize concurrency queue accounting regression #424 修复并合入dev。AdmissionControlTests.CompositeQueueRetryShouldNotConsumeAnUpstreamRatePermitTwice— 三个 rate-policy variants 已由 Stabilize composite admission queue retry regression #425 修复并合入dev。SharpLinkMultiClusterClientTests.DynamicRegistrationShouldRejectASlotChangedWhileItsManifestLoads— 已由 Stabilize manifest-load registration race regression #426 修复并合入dev。AdmissionDynamicRateLegacyWaiterRegressionTests.OldFixedWindowWaiterGrantShouldRemainDebtOnFastTokenBucketTarget— test-only scheduler guard;已由 test(server): remove scheduler bound after legacy rate grant #427 修复并合入dev(squash88ac5b4426ec0665b8b926f972e4be2c1249b7d1)。AdmissionDynamicRateLineageAndLifecycleTests.StopShouldCancelQueuedRateWaiterAndDrainRetiredTimerStateExactlyOnce— test-only ThreadPool continuation latency;已由 test(server): remove scheduler bounds from admission stop drain #428 修复并合入dev(squashfa304c2f8cff806a201c2d82ea8f416f41a4558e)。AdmissionDynamicPartitionUpdateTests.PartitionConcurrencyIncreaseShouldPreserveHolderAndWakeQueuedRequest— post-resize scheduler latency;已由 test(server): remove scheduler bound after partition resize #429 修复并合入dev(squashde458a940473d0bc6e8d42aa9bf2d2d0261fc45a)。UnsizedStreamingPreCreditTests.CreditStarvationShouldBoundLongLivedSerializedOwnersAndWaiters— terminal cleanup 后的本地 scheduler guard;已由 test(runtime): remove scheduler bound after pre-credit terminal #432 修复并合入dev(squashcf8d201cd54ec8fe5bfb990a098a781793cd7db6)。AdmissionDynamicPartitionRateTransitionTests.LateOldPartitionFixedWindowGrantShouldRemainDebtOnTokenBucketTarget— manual-time grant 后的本地 scheduler guard;已由 test(server): remove scheduler bound after old partition rate grant #433 修复并合入dev(squashde985c261b351a5b884d3cd0aa1e5b348225adb7)。AdmissionDynamicPartitionUpdateTests.SelectorReplacementShouldKeepOldQueuedRequestOnOldNamespace— old-namespace holder release 后的本地 scheduler guard;已由 test(server): remove scheduler bound after old selector release #434 修复并合入dev(squashb749bf5c1e045c7c80bc7d80a1b84829e82db2df)。RpcChannelCallShapeIntegrationTests.GeneratedProxyCallsShouldWorkWithRealRpcService(False)—ClientStreamNoReturn totals的一秒 system-clock eventual-completion window;已由 test(integration): remove wall-clock bound from oneway completion poll #436 修复并合入dev(squash7af9a9ab7abd7dcd5d12bebbab99eb47961df7cb)。SharpLinkClientRetryTests.HugeBuiltInJitteredRetryDelayShouldRemainCancellable— 原 flake 来自 retry backoff 前用Task.Delay(20)猜测阶段完成;test(client): synchronize huge retry backoff before stop #439 改用ManualTimeProvidertimer-count barrier,并显式DisableRequestTimeout()排除默认 30s request-timeout timer 的假阳性。32 次 jitter 迭代、两个 2s cancellation/Stop correctness bounds 与ConnectionClosed断言均保留;PR Fast Unit Tests 通过,PR Quick #2613 的唯一失败为无关 transport cleanup case。已由 test(client): synchronize huge retry backoff before stop #439 修复并合入dev(squash8238f93d86016af2258da9ef9ce7429a9888e0e1)。SendPumpProgressIsolationTests.ProgressInterleaveServesProgressWhileNormalQueueNeverEmpties— 原断言把并发产生的 10 个 progress Ping 要求为一个全局连续 wire batch;ci: implement bounded PR Fast gate #437/ci: stop PR Quick from running on every pull request #445 已分别观察到合法 split(例如0..7,72,73与0..3,68..73)。生产契约是 bounded-priority/fairness service,不承诺并发到达的 progress frames 原子成批。test(runtime): assert send-pump progress fairness contract #444 仅删除 contiguous-batch 假设,保留 10 Ping / 130 Response 精确计数以及所有 progress 必须在最后一个 normal frame 前得到服务的 anti-starvation 断言;PR Fast 与 PR Quick Unit Tests 均通过。已由 test(runtime): assert send-pump progress fairness contract #444 修复并合入dev(squash8dc2cd028225a7911b132f8c6084699482eb14e6)。TransportCleanupTests.StreamConnectionDisposeShouldWaitForOutstandingReadRelease— test(client): synchronize huge retry backoff before stop #439 PR Quick #2613 的失败发生在测试已经通过reader.AdvanceTo(...)确定性释放 outstandingReadResult之后,第二个本地DisposeAsync().WaitAsync(2s)guard 超时。ReadOwnershipPipeReader在 completion 请求时同步建立 release TCS,AdvanceTo会确定性完成该 TCS;后续只剩Task.Yield、异步 continuation 与底层PipeReader.CompleteAsync的调度/完成路径,因此该 2s wall-clock bound 不属于 ownership correctness contract。test(runtime): remove scheduler bound after read release #446 仅移除释放后的本地 scheduler bound,仍保留“持有 ReadResult 时 Dispose 不得完成”的核心断言,并在释放后直接等待 disposal 最终完成且验证 owned stream 已关闭;真实 lost wakeup/deadlock 仍会由整体测试超时暴露。PR Fast(含 Unit/Generator/Load Test Tests)通过,已由 test(runtime): remove scheduler bound after read release #446 修复并合入dev(squashf8e42278155a89d4f7a366aa0472d6a32c5f02d4)。SharpLinkServerHostedServiceTests.UnexpectedSuccessfulRunCompletionShouldStopTheHost— ci: stop PR Quick from running on every pull request #445 PR Fast 的原失败来自server.StopAsync(TimeSpan.Zero)后用Task.WhenAny(lifetime.StopRequested.Task, Task.Delay(500ms))猜测 hosted-service observer continuation 能在 500ms 内调度。TestHostApplicationLifetime.StopRequested只由StopApplication()完成,而生产ObserveRunTaskAsync在意外成功退出时正是调用applicationLifetime.StopApplication();test(hosting): await deterministic host stop signal #448 改为直接 await 这个确定性信号,保留并更直接表达“unexpected successful run-loop exit 必须停止 Host”的契约。最终 headd00a4a1bfd6049dde5a12c42d07b715af9d429fc上 PR Fast Bump Microsoft.CodeAnalysis.CSharp to 5.6.0 #46 的 Formatting / maintainability / Release build / dependency / Unit / Generator / Load Test Tests 全绿,CodeQL #2359 也通过。已由 test(hosting): await deterministic host stop signal #448 修复并合入dev(squashbaa4c0ce8722acba24fc6b0651cec395d68567fd)。NegotiatedSessionOptionsTests.DrainingShouldRejectNewRequestsAndPreserveExistingCallFrames— ci: stop PR Quick from running on every pull request #445 PR Fast 的原失败发生在测试已进入 Draining、拒绝新 Request 并排入现有 call frames 之后,额外的FlushSendQueueAsync().AsTask().WaitAsync(2s)wall-clock guard 超时。FlushSendQueueAsync本身会排入 internal force-flush marker 并等待其 completion;send pump 仅在 transportFlushAsync返回并ReleaseBatch后完成该 marker,因此它已经是语义上的 flush barrier。test(runtime): remove scheduler bound from draining flush #449 仅移除 barrier 外层的 2s scheduler bound,所有 draining phase、Unavailablerejection、Response/StreamData/WindowUpdate/Cancel eligibility 断言均保留;真实 flush deadlock/noncompletion 仍会由测试整体超时暴露。最终 head310f5f6ad8b17d0ca1b87c520033f88babd8b1f9上 PR Fast Root referenced service manifests before server build #53(含 Unit/Generator/Load Test Tests)与 CodeQL #2369 全绿。已由 test(runtime): remove scheduler bound from draining flush #449 修复并合入dev(squashfcf28ec4ece1e09337832fd07e39bf5791878963)。RpcSessionLifecycleTests.PumpCreationObservingTerminalShouldReturnValidatedPacket(sync)— 与下一条共用同一 test-infrastructure root cause。test(hosting): await deterministic host stop signal #448 旧-base PR Fast run33300464765中约 2.09s timeout 停在packet.Entered.WaitAsync(2s);StartSendAsync通过LongRunningTestWorker.RunAsync<TResult>启动 dedicated LongRunning worker,但该 overload 原先在Task.Factory.StartNew后立即返回,未证明 worker 已开始执行,因此 runner 调度延迟可能在 packet validation 尚未开始前耗掉 semantic phase budget。test: wait for async long-running worker startup #450 为RunAsync<TResult>补上与 test(client): start lifecycle worker before phase timeout #418 已验证的Run(Action)相同 worker-start barrier;packet validation、terminal winner、validated packet exactly-once return、queue accounting、transport disposal 与 shutdown assertions 均未改。最终 headaa19fa6d7da3464557b13fcb73bfbdd49081ccff上 PR Fast release: stabilize 1.1.1 GoAway gate #58(含 Unit/Generator/Load Test Tests)与 CodeQL #2375 全绿。已由 test: wait for async long-running worker startup #450 修复并合入dev(squash2cbd4e0b9eed5b1eb7ea133ac7910d822127e102)。RpcSessionLifecycleTests.ExistingPumpShouldRejectValidatedPacketAfterTerminalWins(sync)— 同一 test(hosting): await deterministic host stop signal #448 run 中约 2.12s timeout,同样停在packet.Entered.WaitAsync(2s),并经同一StartSendAsync -> LongRunningTestWorker.RunAsync<TResult>路径。test: wait for async long-running worker startup #450 同一 worker-start barrier 修复覆盖该 case;terminal publication、winner identity、exactly-once packet return、queue accounting 与 transport disposal assertions 全部保留。已由 test: wait for async long-running worker startup #450 修复并合入dev(squash2cbd4e0b9eed5b1eb7ea133ac7910d822127e102)。待归因 / 待修复
rpc-add-sharedmemory-c1— allocation spread 超限仍未解决;本轮调查结束,暂不修复。见调查结论。rpc-add-sharedmemory-c8— allocation spread 超限仍未解决;本轮调查结束,暂不修复。见同一调查结论。生产实现、分配预算及正式门禁判定保持原样;停止调查不代表历史失败已被证明无害或已修复。本 tracker 继续保持开放。
后续新增格式
发现新的偶发 CI failure 时继续在本 issue 追加:
处理原则
Thread.Sleep/Task.Delay作为 race 的主要控制手段。完成标准
该 tracker 可以长期保持开放;单个 case 在满足以下条件后标记完成:
Refs #275