Skip to content

[test][ci][tracking] Track newly observed intermittent CI failures #387

Description

@SunSi12138

背景

用于持续记录和归因当前 dev / PR CI 中新出现的偶发测试失败,方便后续集中复现、定位并修复。

这不是 blanket retry / quarantine 列表。每个 case 最终都应尽量判断为:

  • test synchronization / timing issue;
  • runner / platform-sensitive behavior;
  • production concurrency / lifecycle bug;
  • 或已由其他改动修复、不再复现。

历史上一批 timing / runner-sensitive failures 见 #275;本 issue 从当前新观察到的 case 开始继续跟踪,后续可以直接追加新条目。

当前状态

已处理 / 已合并

待归因 / 待修复

  • rpc-add-sharedmemory-c1 — allocation spread 超限仍未解决;本轮调查结束,暂不修复。见调查结论。
  • rpc-add-sharedmemory-c8 — allocation spread 超限仍未解决;本轮调查结束,暂不修复。见同一调查结论。

生产实现、分配预算及正式门禁判定保持原样;停止调查不代表历史失败已被证明无害或已修复。本 tracker 继续保持开放。

后续新增格式

发现新的偶发 CI failure 时继续在本 issue 追加:

### N. `TestName`

现象:
证据 / run:
环境:
初步关注点:
状态: 待复现 / 已归因 / 已修复 / 不再复现
关联 PR / issue:

处理原则

  1. 先记录失败 run / commit / runner / exception 或 timeout 证据,再做归因。
  2. focused repetition / stress 优先于反复 rerun 整个 CI。
  3. 对 concurrency / lifecycle race,优先增加 deterministic synchronization seam。
  4. 不用 Thread.Sleep / Task.Delay 作为 race 的主要控制手段。
  5. 不以 blanket retry、skip、quarantine 或单纯扩大 timeout 作为最终修复。
  6. 如果确认是独立 production bug,可以拆出单独 issue / PR,并在这里回链。
  7. 修复后记录 root cause、验证方式以及是否完成重复运行验证。

完成标准

该 tracker 可以长期保持开放;单个 case 在满足以下条件后标记完成:

  • root cause 已明确;
  • 修复或合理处置已落地并回链;
  • focused 重复验证不再复现;
  • 相关 Unit / Integration / CI gate 通过;
  • 没有通过弱化 correctness assertion 来“稳定”测试。

Refs #275

Activity

  1. SunSi12138 commented on Aug 29, 2026

    @SunSi12138
    OwnerAuthor

    Client-stream time-budget tests can race request registration / emission

    Observed in PR #386, PR Quick #2410 at 97059f14443b167ced62db0c67eedbe5ac4a2fe2.

    Failures:

    • TimedOneWayClientStreamShouldNotStartProducerUntilRequestSurvivesEmission: expected local SharpLinkException, but the invocation completed without one.
    • TimedClientStreamShouldNotStartProducerUntilRequestSurvivesEmission: timed out after 30s with the invocation still WaitingForActivation.

    Classification: pre-existing test synchronization race; not introduced by #386.

    Evidence:

    • Both failing test files are byte-identical between dev (00e2f18c6384c785d232bd59902102d3af7ad3da) and PR feat: assembly-owned codec routing (clean port of #320) #386 head.
    • The previous PR head f0bd386a8534aff77ed1086c334695f90f8bde9c passed all 1364 unit tests under the same Ubuntu runner image / .NET SDK.
    • The only delta from that green head to 97059f... is a generator-only commit; it does not change request deadlines, the send queue, or send-pump timing.
    • In TimedClientStreamShouldNotStartProducerUntilRequestSurvivesEmission, the test starts the async invocation and then immediately advances ManualTimeProvider and explicitly flushes. There is no deterministic barrier proving that the invocation has already registered its 5s deadline and enqueued the initial Request. If the test thread wins that race, the explicit flush can observe an empty queue; the Request is then created against the already-advanced clock and can remain behind the 10s manual-time session flush timer, matching the observed 30s wall-clock test timeout.
    • The one-way variant also injects clock advancement through a transport-side “next output buffer request” hook. Its failure shape is consistent with the same missing target-request synchronization around the emission boundary.

    Fix direction: add deterministic synchronization tied to the target Request reaching the enqueue/emission boundary. Do not add retries or increase timeouts. I’m preparing the follow-up fix separately from #386.

  2. SunSi12138 commented on Aug 29, 2026

    @SunSi12138
    OwnerAuthor

    Follow-up fix #392 has been merged into dev as 3821a59f307351e6c41c7675b4ff02a2da9de005.

    The rerun of #386's original failed PR Quick attempt passed both client-stream time-budget tests, confirming the failures were intermittent rather than caused by #386. #392 removes the synchronization race without retries or larger timeouts by draining pre-existing connection output and advancing the manual clock at the target Request's output-buffer acquisition.

    Validation before merge: formatting, Debug/Release build, and Unit Tests passed.

  3. SunSi12138 commented on Aug 29, 2026

    @SunSi12138
    OwnerAuthor

    New intermittent evidence observed while validating #386 on exact head df5526e322f4bfb947909c521d4e4612577a94f0 in PR Quick #2422:

    • attempt 1 quick Unit Tests: 1363/1365 passed; two existing tests hit their local ~2s timing guards:
      • PooledAsyncStreamDispatcherLocalAbortTests.LocalAbortShouldWaitForOwnedBufferedPublication
      • SharpLinkClientLifecycleStateTests.StopAsyncShouldNotRunShutdownCallbacksBeforeReturning
    • The test source blobs are unchanged from dev@3821a59f307351e6c41c7675b4ff02a2da9de005 (f87b573... and 6ea9027... respectively).
    • No production/test change was made for either failure. A GitHub failed-job rerun on the same head then completed Unit Tests 1365/1365 successfully.
    • The feat: assembly-owned codec routing (clean port of #320) #386 changes do not alter the StopAsync lifecycle body; its client lifecycle diff is codec-provider construction plumbing. The LocalAbort case directly exercises the dispatcher and does not traverse the new request-stream drain API.

    Current classification: scheduling-sensitive / intermittent evidence, not yet a proven #386 production regression. Per this tracker policy, do not widen timeouts or add blanket retries. If either case repeats, run focused Linux repetition/stress and separate test-barrier stabilization from production changes based on the reproduction.

  4. SunSi12138 commented on Aug 29, 2026

    @SunSi12138
    OwnerAuthor

    New intermittent evidence while validating #386 on exact head b1e0df8f477f699bdfb2b16c9a796a5d72519006, PR Quick #2463 (quick job 99133285778):

    Unit Tests: 1362/1364 passed; two tests hit local ~2s timeout guards.

    1. SharpLinkMultiClusterClientTests.DynamicRegistrationShouldRejectASlotChangedWhileItsManifestLoads

      • SharpLinkMultiClusterClientTests.cs:1113
      • failed after ~2.01s with TimeoutException: The operation has timed out.
      • I could not find an earlier issue/tracker occurrence by test name; treat this as a newly observed intermittent case, not as previously established evidence.
    2. PooledAsyncStreamDispatcherLocalAbortTests.LocalAbortShouldWaitForOwnedBufferedPublication

      • PooledAsyncStreamDispatcherLocalAbortTests.cs:129
      • failed after ~2.04s with TimeoutException: The operation has timed out.
      • this is a repeat of the case already recorded here from PR Quick #2422 on df5526e..., where a same-head failed-job rerun subsequently passed Unit Tests 1365/1365 without production/test changes.

    The #386 cleanup head compiled successfully in Debug and Release with 0 warnings / 0 errors before these Unit failures, and the current #386 changes are generator/routing/identity-boundary cleanup rather than the MultiCluster dynamic-registration or dispatcher timing paths.

    Classification for now: scheduling-sensitive/intermittent evidence, not a proven #386 production regression. Keep these cases in #387; do not widen timeouts or add blanket retries. If DynamicRegistrationShouldRejectASlotChangedWhileItsManifestLoads repeats, run focused Linux repetition/stress and determine whether it needs its own production/test synchronization issue.

  5. SunSi12138 commented on Aug 29, 2026

    @SunSi12138
    OwnerAuthor

    Admission-control queue timing/scheduling failures during #386 validation

    Observed on exact #386 head 95a03ff8b53abca7ccfcb90e0c3347abc11eb807, PR Quick #2470 (quick job), Ubuntu 24.04 / .NET SDK 10.0.400.

    Unit Tests: 1359/1364 passed; five admission-control/queue cases failed:

    • AdmissionControlTests.ConcurrencyQueueShouldReleasePermitAndAccountingExactlyOnce — ~3.62s; assertion queued call acquired.
    • AdmissionControlTests.CompositeQueueRetryShouldNotConsumeAnUpstreamRatePermitTwice(fixed) — ~3.57s; assertion queued request should reuse its previously consumed rate permit.
    • AdmissionStateKernelTests.OverlappingGenerationsShouldShareQueueBoundsAndRetainedBytes(2, 3, queue_bytes) — ~2.51s TimeoutException.
    • AdmissionStateKernelTests.OverlappingGenerationsShouldShareQueueBoundsAndRetainedBytes(1, 8, queue_count) — ~2.48s TimeoutException.
    • AdmissionControlTests.QueueTimeoutShouldReleasePartitionOwnershipExactlyOnce — TUnit 30s timeout, task remained WaitingForActivation.

    I found no prior #387 entry for these names. #386 does not change the AdmissionControl/AdmissionStateKernel test files or admission-control implementation; the current work is codec generator/routing ownership cleanup. Treat this as newly observed scheduling/intermittent evidence, not as a #386 codec regression unless focused reproduction proves otherwise.

    Per tracker policy: do not widen timeouts or add blanket retries. Re-run the same #386 head to unblock Generator Tests; if these recur, investigate them separately with deterministic queue/permit synchronization.

  6. SunSi12138 commented on Aug 30, 2026

    @SunSi12138
    OwnerAuthor

    Admission dynamic waiter/resize timing failures observed on PR #420

    Observed while validating PR #420 on exact head 7c2cf51f8b299eca3112ac863edf4d87517fd496, PR Quick #2499 (quick job 99207503729), Ubuntu 24.04 / .NET SDK 10.0.400.

    Unit Tests: 1360/1363 passed; three Admission dynamic-update cases hit local timeout guards:

    1. AdmissionDynamicRateLineageAndLifecycleTests.StopShouldCancelQueuedRateWaiterAndDrainRetiredTimerStateExactlyOnce

      • failed after ~2.17s with TimeoutException
      • the test is waiting for a queued rate waiter to complete after AdmissionStateKernel.DisposeAsync()/Stop begins draining, then verifies queue accounting, retired state, and timer disposal exactly once.
    2. AdmissionDynamicRateLegacyWaiterRegressionTests.OldFixedWindowWaiterGrantShouldRemainDebtOnFastTokenBucketTarget

      • failed after ~2.33s with TimeoutException
      • after replacing an old FixedWindow generation with a fast TokenBucket target and advancing manual time by 40s, the old captured-generation waiter did not complete within its 2s wall-clock guard.
    3. AdmissionDynamicPartitionUpdateTests.PartitionConcurrencyIncreaseShouldPreserveHolderAndWakeQueuedRequest

      • failed after ~7.18s with TimeoutException
      • the test updates partition concurrency 1 -> 2 and expects the existing FIFO waiter to wake and acquire the newly available permit; the waiter did not complete in the expected window.

    PR #420 changes only documentation / project-reference policy files and does not modify Admission production or test code, so these failures are not currently evidence of a #420 regression.

    This is the same broad failure family as the Admission-control queue timing/scheduling cases already recorded in this tracker from PR Quick #2470, but these three exact test names were not previously listed.

    Current classification: new intermittent / scheduling-sensitive Admission waiter evidence; root cause not yet established. Do not widen timeout guards or add blanket retries. If any of these repeat, run focused Linux repetition/stress with deterministic queue/timer/resize synchronization and determine whether the wakeup path exposes a production lifecycle bug or only test orchestration sensitivity.

  7. SunSi12138 commented on Aug 30, 2026

    @SunSi12138
    OwnerAuthor

    Tracker update: each remaining independent #387 root cause now has its own non-draft PR against dev (all Ready for review):

    The two client-stream timing cases already fixed by merged #392 were intentionally not duplicated.

    #423 needed one follow-up after CI showed the manual timer callback is allowed to complete asynchronously; the updated test now removes only the unrelated wall-clock completion bound, and its latest Unit Tests pass. #424 also received a formatting-only follow-up for the required final newline. #425 is fully green; #426 has passed formatting/build/unit/integration/native-AOT portions while the workflow finishes its remaining smoke/report tail.

    Keeping #387 open until the child PRs are reviewed/merged.

  8. SunSi12138 commented on Aug 30, 2026

    @SunSi12138
    OwnerAuthor

    Additional unrelated intermittent failures observed while reviewing #416 / #421 / #424

    While reviewing the #387 stabilization PRs, several CI failures appeared on PRs whose changes do not touch the failing path. Some related Admission cases were already tracked here, but the following exact test names were not previously recorded.

    1. RpcChannelCallShapeIntegrationTests.GeneratedProxyCallsShouldWorkWithRealRpcService(False)

    Observed on PR #416 head 158763c4d32ea70d9375d5630fa67b0f0b4c3c24, PR Quick #2518 (33293669721), Ubuntu 24.04 / .NET SDK 10.0.400.

    • Integration Tests: 409/410 passed.
    • Failure: assertion ClientStreamNoReturn totals in RpcChannelCallShapeIntegrationTests.cs:166, reached from the generated-proxy real-service test around line 91.
    • test(runtime): synchronize local-abort publication ownership #416 changes only PooledAsyncStreamDispatcherLocalAbortTests test synchronization; it does not modify RPC generated-proxy call-shape production or integration-test code.

    Classification: new intermittent integration evidence; unrelated to #416; root cause not yet established. If it repeats, investigate client-stream no-return completion/accounting synchronization rather than widening the eventual-consistency timeout.

    2. AdmissionDynamicPartitionRateTransitionTests.LateOldPartitionFixedWindowGrantShouldRemainDebtOnTokenBucketTarget

    Observed on PR #421 head 5d824a16e10a1009dc2b05bf9ac59dca4041af33, PR Quick #2497 (33292604360).

    • Failed after ~2.35s with local TimeoutException around AdmissionDynamicPartitionRateTransitionTests.cs:175.
    • test(server): remove scheduler bound after admission permit release #421 is test-only and changes only AdmissionStateKernelTests.OverlappingGenerationsShouldShareQueueBoundsAndRetainedBytes; it does not change this test or Admission production code.
    • This is related to the broader Admission generation/waiter timing family already seen in this tracker, but this exact test name had not been listed.

    Classification: new scheduling-sensitive Admission generation/waiter evidence; unrelated to #421.

    3. AdmissionDynamicPartitionUpdateTests.SelectorReplacementShouldKeepOldQueuedRequestOnOldNamespace

    Observed in the same PR #421 / PR Quick #2497 run.

    Classification: new scheduling-sensitive Admission selector/queued-waiter evidence; unrelated to #421; root cause not yet established.

    Note: the same run also failed ConcurrencyQueueShouldReleasePermitAndAccountingExactlyOnce, but that exact case was already recorded in #387 and was subsequently addressed by #424, so it is not duplicated here.

    4. UnsizedStreamingPreCreditTests.CreditStarvationShouldBoundLongLivedSerializedOwnersAndWaiters

    Observed on PR #424 head f1e1ca08111d311633ca3c7d86beb1eeef157659, PR Quick #2514 (33293525861).

    • Failed after ~2.11s with local TimeoutException in ExpectSameException, around UnsizedStreamingPreCreditTests.cs:153 / test line 84.
    • Stabilize concurrency queue accounting regression #424 is test-only and changes one Admission concurrency-queue test to use ManualTimeProvider; it does not touch streaming pre-credit production or tests.

    Classification: new scheduling-sensitive streaming pre-credit evidence; unrelated to #424; root cause not yet established.

    Note: the same run also failed the three CompositeQueueRetryShouldNotConsumeAnUpstreamRatePermitTwice variants; that exact family was already recorded here and addressed by #425, so it is not duplicated.

    Per #387 policy: keep these as evidence, do not add blanket retries or simply enlarge wall-clock guards. If they recur, use focused Linux repetition/stress and deterministic phase barriers to separate test orchestration sensitivity from a production concurrency/lifecycle defect.

  9. SunSi12138 commented on Aug 30, 2026

    @SunSi12138
    OwnerAuthor

    Follow-up on the three Admission dynamic waiter/resize failures from PR Quick #2499:

    • test(server): remove scheduler bound after legacy rate grant #427 — AdmissionDynamicRateLegacyWaiterRegressionTests.OldFixedWindowWaiterGrantShouldRemainDebtOnFastTokenBucketTarget
      • root cause classified as a test-only scheduler guard after ManualTimeProvider.Advance(40s) has already triggered the source-generation grant
      • removes only the unrelated 2s wall-clock WaitAsync; transition-debt/expiry assertions are unchanged
    • test(server): remove scheduler bounds from admission stop drain #428 — AdmissionDynamicRateLineageAndLifecycleTests.StopShouldCancelQueuedRateWaiterAndDrainRetiredTimerStateExactlyOnce
      • root cause classified as test-only ThreadPool continuation latency after synchronous kernel draining cancellation
      • removes the two local 2s scheduler guards while preserving shutdown error, queue accounting, state reclamation, and timer-disposal assertions
    • test(server): remove scheduler bound after partition resize #429 — AdmissionDynamicPartitionUpdateTests.PartitionConcurrencyIncreaseShouldPreserveHolderAndWakeQueuedRequest
      • root cause classified as test-only scheduler latency after the synchronous 1 -> 2 partition concurrency target commit has woken the queued waiter
      • removes only the post-resize 2s wall-clock guard; namespace/holder/FIFO/accounting assertions are unchanged

    Each case is isolated in its own non-draft PR against dev; CI is running.

    #412 was also revised after review found a race in its first fix: the separate _writerActive == 0 / pending-empty observation could be invalidated by a concurrent outbound signal before Close enqueue, allowing an earlier blocked write to bypass the 250ms cleanup path. The updated #412 serializes writer active state, pending enqueue/dequeue, and Dispose's idle + empty -> Close transition under one outbound-state gate. Unbounded waiting for Close-start is now used only when Dispose atomically owns a genuinely idle/empty writer; otherwise the original bounded cleanup path is preserved.

  10. SunSi12138 commented on Aug 30, 2026

    @SunSi12138
    OwnerAuthor

    Additional unrelated intermittent failure observed while reviewing #433

    SharpLinkClientRetryTests.HugeBuiltInJitteredRetryDelayShouldRemainCancellable

    Observed on PR #433 head b8f493a927b0e935831707a51e052f656953d5c7, PR Quick #2553 (33295641662), Ubuntu 24.04 / .NET SDK 10.0.400.

    • Unit Tests: 1362/1363 passed.
    • Failure: local TimeoutException after ~4.22s at SharpLinkClientRetryTests.cs:173-174.
    • test(server): remove scheduler bound after old partition rate grant #433 changes only AdmissionDynamicPartitionRateTransitionTests.LateOldPartitionFixedWindowGrantShouldRemainDebtOnTokenBucketTarget; it does not touch client retry/jitter/cancellation production or test code.
    • Formatting, maintainability gates, Debug/Release builds, generated-assembly dependency checks, CodeQL, and codec-compatibility jobs passed before/alongside the unrelated unit failure.

    Classification: new scheduling/timing-sensitive client retry cancellation evidence; unrelated to #433; root cause not yet established.

    Per #387 policy, keep this as evidence rather than widening the local timeout or adding blanket retries. If it repeats, use focused repetition with a deterministic barrier around cancellation registration / retry-delay scheduling.

  11. SunSi12138 commented on Aug 30, 2026

    @SunSi12138
    OwnerAuthor

    Tracker follow-up for the two currently unchecked cases in #387:

    • test(integration): remove wall-clock bound from oneway completion poll #436 — RpcChannelCallShapeIntegrationTests.GeneratedProxyCallsShouldWorkWithRealRpcService(False) / ClientStreamNoReturn totals

      • root cause: the exact total includes [Oneway] client-stream handlers, whose server-side completion is not implied by client-side return; the helper imposed an unrelated 1s system-clock deadline on eventual server completion
      • fix keeps the exact total assertion and polling, removes only that local wall-clock deadline
      • non-draft / Ready for review; PR Quick and CodeQL are green
    • test(client): synchronize huge retry backoff before stop #439 — SharpLinkClientRetryTests.HugeBuiltInJitteredRetryDelayShouldRemainCancellable

      • root cause: Task.Delay(20) guessed that the built-in huge jittered retry backoff had already been scheduled before StopAsync
      • fix uses ManualTimeProvider timer registration as a deterministic phase barrier before Stop, while retaining all 32 iterations, the original 2s cancellation/Stop bounds, and the expected ConnectionClosed result
      • non-draft / Ready for review; PR Quick and CodeQL are green

    Both fixes are test-only; neither changes production RPC/retry/timeout semantics. Keeping #387 open for review/merge and future intermittent evidence.

  12. SunSi12138 commented on Aug 30, 2026

    @SunSi12138
    OwnerAuthor

    SendPumpProgressIsolationTests.ProgressInterleaveServesProgressWhileNormalQueueNeverEmpties

    现象:

    证据 / run:

    环境:

    • Ubuntu 24.04 hosted runner
    • .NET SDK 10.0.400
    • PR Quick 失败时 Unit Tests: 1362 passed / 1 failed

    初步关注点:

    • 该测试依赖“progress frames 必须作为一个连续 batch 出现在 wire order 中”的调度假设;本次实际观察到 8 个 Ping 在开头、剩余 2 个 Ping 在 index 72/73,说明 pump/pipe 调度可以把这 10 个 progress frame 分成多个合法可见批次。
    • 历史 test(runtime): stabilize progress starvation regression #291 曾修过同一 SendPumpProgressIsolationTests 文件中的另一个 scheduler-dependent starvation regression test,但不是这个具体 case;本条按新 flake 单独记录。
    • 按 tracker 原则,优先做 focused Linux repetition/stress,并确认该 contiguous-batch assertion 是否真的是 production contract;不要通过 blanket retry、扩大 timeout 或简单弱化 correctness assertion 处理。

    状态: 待复现 / 待归因
    关联 PR / issue: #437, #291

  13. SunSi12138 commented on Aug 30, 2026

    @SunSi12138
    OwnerAuthor

    New intermittent evidence from #439 / PR Quick #2613

    Observed on #439 head 0d86048f60aa20fcb56e138c88bf6ba4ebede139, PR Quick #2613 (33298940958), Ubuntu 24.04.4 / .NET SDK 10.0.400.

    Failure:

    • TransportCleanupTests.StreamConnectionDisposeShouldWaitForOutstandingReadRelease
    • Unit Tests: 1362/1363 passed; this was the only failure.
    • ~2.781s TimeoutException at TransportCleanupTests.cs:44.

    The timeout is the second local 2s guard: after disposal has requested reader completion, the test verifies disposal is still pending while the consumer owns the ReadResult, calls reader.AdvanceTo(result.Buffer.End) to release that ownership, and then times out waiting for connection.DisposeAsync() to finish.

    Classification: new intermittent / lifecycle-sensitive evidence; unrelated to #439, root cause not yet proven test-only.

    Evidence for unrelatedness:

    The production path has deterministic read-release signaling, but completion after release still crosses asynchronous continuations (RunContinuationsAsynchronously / Task.Yield) and the underlying PipeReader.CompleteAsync. One CI occurrence is not enough to decide whether the local 2s guard is merely scheduler-sensitive or whether there is a latent lifecycle/completion race.

    Next step should be focused Linux repetition/stress of this exact test and inspection of the post-AdvanceTo completion path. Do not widen the timeout or add blanket retries as the fix.

  14. SunSi12138 commented on Aug 30, 2026

    @SunSi12138
    OwnerAuthor

    SharpLinkServerHostedServiceTests.UnexpectedSuccessfulRunCompletionShouldStopTheHost

    Observed in PR #445, PR Fast run 33299359487, merge commit 477142fd1e072a41012d7606b2036a00599f5edb (head ae8c615239f830cc11028b395ecda967373756da).

    现象:

    • Unit Tests: 1362/1363 passed; only this test failed.
    • Failure assertion: an unexpected successful Server run-loop exit must stop the owning Host.
    • Test elapsed ~608ms and uses Task.WhenAny(lifetime.StopRequested.Task, Task.Delay(500ms)) as a wall-clock guard after calling server.StopAsync(TimeSpan.Zero).

    证据 / run:

    • ci: stop PR Quick from running on every pull request #445 changes exactly one workflow line: removes pull_request from .github/workflows/pr-quick.yml; it does not modify Hosting, Server, or test code.
    • Restore, formatting, maintainability gates, Release build, and generated-assembly dependency verification all passed before Unit Tests.
    • SharpLinkServerHostedService.ObserveRunTaskAsync observes the server run task asynchronously and calls applicationLifetime.StopApplication() when it sees an unexpected successful completion. The test has no deterministic barrier for that continuation; it only waits up to 500ms wall-clock.
    • Runner: Ubuntu 24.04, .NET SDK 10.0.400.

    初步关注点:

    • Strong candidate for test synchronization / scheduler-latency flake rather than a ci: stop PR Quick from running on every pull request #445 regression.
    • The 500ms wall-clock Task.Delay guard can lose under hosted-runner / parallel-test scheduling pressure even when the production continuation is correct.
    • Do not fix by merely widening the timeout or adding blanket retry. Prefer a deterministic observation/barrier around StopApplication() / run-task completion if reproduction confirms this classification.

    状态: 待复现 / 待归因。A rerun of the exact failed Fast job has been started; record the rerun result here once complete.

    关联 PR: #445

  15. SunSi12138 commented on Aug 30, 2026

    @SunSi12138
    OwnerAuthor

    New PR Fast unit-test flakes observed on #445

    Observed in PR #445, PR Fast run 33299548254 on merge commit 460c9369994bdb0d82400d9b7a937f940b44f2e3 (head 6caa29679b63f1cc78d123b68b80edea424b314f, base 8238f93d86016af2258da9ef9ce7429a9888e0e1), Ubuntu 24.04 / .NET SDK 10.0.400.

    Unit Tests result: 1361/1363 passed; two failures:

    1. SendPumpProgressIsolationTests.ProgressInterleaveServesProgressWhileNormalQueueNeverEmpties

      • Failed in 43ms.
      • Assertion expected all 10 progress Ping frames to form one contiguous batch.
      • Observed Ping indices: 0,1,2,3,68,69,70,71,72,73.
      • The test enqueues 130 normal frames and then sequentially enqueues 10 Ping frames with no deterministic barrier preventing the send pump from waking after only part of the Ping producer loop has completed.
      • Production send-pump policy drains progress at the loop top and every 64 normal frames. The observed split is consistent with a legal schedule: the pump drains the first 4 Pings, processes 64 normal frames, then drains the remaining 6 at the interleave boundary.
      • The test file is byte-identical on ci: stop PR Quick from running on every pull request #445 base/head (6a0cdfd79df65768e8ac9a141d5f81b5c954e865). ci: stop PR Quick from running on every pull request #445 itself only changes the PR Quick workflow trigger.
      • Initial classification: test synchronization / scheduling assumption, not evidence of a ci: stop PR Quick from running on every pull request #445 production regression. The intended invariant should be clarified: bounded progress service is guaranteed, but atomic contiguity of concurrently-enqueued progress frames is not obviously a runtime contract.
    2. NegotiatedSessionOptionsTests.DrainingShouldRejectNewRequestsAndPreserveExistingCallFrames

      • Failed with TimeoutException after ~3.27s at FlushSendQueueAsync().AsTask().WaitAsync(TimeSpan.FromSeconds(2)).
      • Test source is byte-identical on ci: stop PR Quick from running on every pull request #445 base/head (4b5c0e2cf4721658f82bb3626ab66f9a527028c6).
      • Initial classification: scheduler/lifecycle-sensitive evidence; root cause not yet established. Focused repetition should determine whether this is send-pump flush continuation latency or a real draining/flush race.

    Both failures are unrelated to #445's one-line workflow-trigger diff. Per #387 policy, do not fix either by merely increasing timeouts, blanket retry, or weakening the correctness assertion. Prefer deterministic synchronization / focused stress and confirm the actual production invariant first.

  16. 79 remaining items

  17. added a commit that references this issue on Sep 18, 2026
  18. SunSi12138 commented on Sep 18, 2026

    @SunSi12138
    OwnerAuthor

    2026-09-18 update: Nightly Regression #284/#285 follow-up has been resolved and merged into dev.

    • fix: close deadline arm race and stabilize #387 lifecycle regressions #698 (e355e0f3101756f94fbae8f0ee45ad0bb22061d9, squash) fixes the production absolute-deadline timer arm TOCTOU and adds deterministic full/partial arm-drift plus caller-cancellation arbitration regressions. It also replaces the disconnected-cleanup test's premature ownership assertion with the framework-supervisor completion barrier, and removes the per-round 2s remote stream-disposal wall-clock guard while retaining the enclosing test timeout.
    • test(client): synchronize dynamic reconnect replacement timer #699 (e1edc26b292ee0c464f0f94aa31895c45e93b21c, squash) fixes the macOS dynamic reconnect replacement flake by waiting for the exact live 1s reconnect timer before advancing manual time; production reconnect semantics are unchanged.

    Both PRs were reviewed with no remaining correctness blockers and their required fast checks passed before merge. The tracker remains open for future intermittent CI failures.

  19. SunSi12138 commented on Sep 18, 2026

    @SunSi12138
    OwnerAuthor

    2026-09-18 follow-up: #701 has been merged into dev as a2d1ee3b6e330646fc2677bdf489cddcfdcd6698.

    The Browser codec evidence policy is now narrowed to the documented compatibility boundary: only Browser <-> hosted-desktop failures in the existing auto-layout-release-scoped category are retained as non-blocking evidence when they have one of the recognized layout/decode mismatch shapes. This covers the known Mono/wasm32 vs CoreCLR/64-bit managed-layout variance that was blocking the main Release Gate.

    Browser self-roundtrip, desktop <-> desktop compatibility, every other fixture category, fixture/identity/hash/completeness checks, and unexpected/malformed classifications remain blocking. No production codec hot-path change, retry, timeout widening, or blanket continue-on-error was introduced.

    #700 has been refreshed to the new dev tree and will rerun the Release Gate.

  20. SunSi12138 commented on Sep 18, 2026

    @SunSi12138
    OwnerAuthor

    2026-09-18 Browser evidence follow-up: #702 merged into dev as 62d876d01321c0afac22b56c54a5b40039180481.

    The first #701 Release Gate rerun confirmed the intended C# policy behavior: the known Browser -> desktop auto-layout rows were emitted with blocking=false and the verifier reported blocking failures: 0. The remaining failure was only the JS strict evidence-shape checker: for DESERIALIZE_REJECTED, Fixture<T>.Verify can retain a previously computed logicalEquality=false when diagnostic value rendering throws.

    #702 narrows that validator mismatch by allowing exactly null or false in that field for the already evidence-only rejected row; true, other classifications, other fixture categories, and other evidence invariants remain rejected. #700 has been refreshed again to the current dev tree.

  21. SunSi12138 commented on Sep 18, 2026

    @SunSi12138
    OwnerAuthor

    2026-09-18 Release Gate #606 Windows integration flake:

    • Run 35336209933, job 105571522074
    • SharedMemoryListenerShouldRejectBadHandshakesAndAcceptNextClient failed after the three bad-handshake cases with Shared-memory transport handshake timed out after 00:00:00.1000000 on the final healthy client.
    • The listener's 100ms handshake timeout starts only after a pipe connection is accepted; the client factory's same 100ms timeout starts before ConnectAsync, so it also charges the scheduler/accept-loop tail while the listener cycles from the prior rejected connection.
    • The server-side 100ms rejection contract remains valid; the flake is caused by reusing that narrow server test budget for the final healthy recovery client's local end-to-end wait.

    A focused test-only follow-up is being prepared to keep the listener at 100ms while giving only the final healthy client its own recovery timeout. No production transport behavior or bad-handshake timeout is being widened.

  22. SunSi12138 commented on Sep 18, 2026

    @SunSi12138
    OwnerAuthor

    2026-09-18 follow-up: #703 and #704 are now merged into dev.

    The Browser CoreCLR graph is intentionally non-blocking while the runtime is pre-GA. Existing Mono Browser evidence and the release-gated desktop contract remain unchanged. #700 has been refreshed to the new dev tree so Release Gate can execute the new evidence lane.

  23. SunSi12138 commented on Sep 18, 2026

    @SunSi12138
    OwnerAuthor

    2026-09-18 Browser CoreCLR identity follow-up: #705 merged into dev as 400f490a4ca18223ee6acf3b5f6c742ee2e4987b.

    Release Gate #607 established that .NET 11 RC1 Browser CoreCLR successfully publishes and runs in headless Chrome with UseMonoRuntime=false. It also exposed that Type.GetType("Mono.Runtime") returns null on the existing Browser/Mono probe, so reflection cannot be used to distinguish the Browser runtime family.

    #705 now records Browser runtime family from the actual project runtime selection: default net10 Browser remains Mono/platform-runtime-pack, while the explicit UseMonoRuntime=false build records CoreCLR/build-runtime-selection. Observed TFM/RID/pointer size/framework/runtime/SDK/compilation mode remain runtime evidence.

    #700 has been refreshed again so the full evidence graph can be rerun with corrected identities.

  24. SunSi12138 commented on Sep 18, 2026

    @SunSi12138
    OwnerAuthor

    2026-09-18 final follow-up: the CI stability / Browser codec work has now landed on main.

    No new tag or GitHub Release was created; the latest published release remains v2.0.0. The tracker remains open for future intermittent CI failures.

  25. SunSi12138 commented on Sep 19, 2026

    @SunSi12138
    OwnerAuthor

    2026-09-19 documentation follow-up: codec compatibility terminology is now aligned with the current CI evidence.

    This is documentation/tracking alignment only; no production or test-gating behavior was changed.

  26. SunSi12138 commented on Sep 19, 2026

    @SunSi12138
    OwnerAuthor

    新增 intermittent CI failure — StaticClusterReconnectShouldBeSingleFlightAtTheProviderBoundary

    现象:

    • PR feat: add optional generation control extension #706 在最终 head c8f8b21219643b5614055a1adaa609eeea960485 上,一轮由 PR description edit 触发的 PR Fast 中 Unit Tests 失败。
    • 1851 个 Unit Tests 中仅此 1 个失败。
    • 失败信息:endpoint static-first has no active reconnect owner。
    • 栈位于 SharpLinkClientLifecycleReconnectSupport.GetStaticReconnectTask(...):71,由 SharpLinkClientLifecycleReconnectTests.StaticClusterReconnectShouldBeSingleFlightAtTheProviderBoundary():150 触发。

    证据 / run:

    • 失败 run: PR Fast 35441391215,首次 fast job 105892663366。
    • 同一 head 的前一轮 PR Fast 35441150436 已完整通过。
    • 对失败 job 原样重跑后,新 fast job 105893110493 成功,Unit Tests、Generator Tests 及后续 Fast checks 全部通过。
    • 同一 head 的 PR Extended 35441347233 也成功,包括 Integration Tests、NativeAOT smoke、package smoke、demo/load 及桌面 codec compatibility。

    环境 / 变更相关性:

    • 该测试文件 test/SharpLink.UnitTests/Client/SharpLinkClientLifecycleReconnectTests.cs 与 helper SharpLinkClientLifecycleReconnectSupport.cs 在 PR head 与当前 dev 上 blob SHA 完全一致。
    • PR feat: add optional generation control extension #706 不修改 Client reconnect production path 或上述测试。
    • 因此当前证据更符合 existing timing / reconnect-owner observation flake,而不是 GenerationControl 回归。

    初步关注点:

    • 测试在读取 StaticClusterRuntime reconnect owner 时,可能存在“owner 已被消费/完成但测试尚未观察到下一 owner”或 provider-boundary single-flight 状态切换的短暂窗口。
    • 若后续复现,优先给 reconnect owner publication / timer arm / ownership handoff 增加确定性 barrier,再做 focused repetition;不要用 blanket retry、sleep 或单纯扩大 timeout 作为最终修复。

    状态: 待复现 / 待归因。

    关联:

  27. SunSi12138 commented on Sep 19, 2026

    @SunSi12138
    OwnerAuthor

    新增 intermittent CI failure — TimedBatchShouldExtendBatchForFrameArrivingBeforeDeadline

    现象:

    • Release Gate run 35448326880 的 matrix-build-test (macos-latest) 唯一 Unit Test 失败。
    • 1851 个 Unit Tests 中仅此 1 个失败。
    • 失败信息:TimeoutException: The operation has timed out.
    • 栈定位到 test/SharpLink.UnitTests/Runtime/SendPumpTests.cs:174 / :188。
    • 失败用时约 2.18s,属于测试内 2s timing guard 超时。

    证据 / run:

    • Release Gate run: 35448326880
    • failed job: 105911035576
    • 环境:macOS arm64 Release Unit Tests
    • 同一 merge commit 的 PR Fast、PR Package Smoke、CodeQL、Codec Padding Security Evidence、Nightly Regression 均成功;Release Gate 其余 Windows/Linux matrix、AOT、chaos、pack、codec lanes 也通过。

    变更相关性:

    初步关注点:

    • 检查测试对 timed-batch deadline / second-frame arrival 的同步是否依赖 wall-clock scheduler timing。
    • 若可复现,优先引入 deterministic synchronization/manual-time seam,而不是扩大 timeout 或 blanket retry。

    状态: 待复现 / 待归因。

    关联:

  28. SunSi12138 commented on Sep 19, 2026

    @SunSi12138
    OwnerAuthor

    Follow-up for TimedBatchShouldExtendBatchForFrameArrivingBeforeDeadline:

    • the exact failed Release Gate macOS job was re-run unchanged
    • re-run job: 105914248807
    • result: success
    • Unit Tests passed, including the previously failing timed-batch case
    • Generator Tests, Load Test Tests, Integration Tests, Pack, and Package Smoke also passed

    No code or timeout threshold was changed. This confirms the observed failure was intermittent on the same code/commit; keep the case tracked for deterministic synchronization/root-cause work rather than treating it as a GenerationControl regression.

  29. SunSi12138 commented on Sep 21, 2026

    @SunSi12138
    OwnerAuthor

    新增 intermittent CI failure — allocation gate stability spread

    现象:

    • PR perf: hoist streaming codec capability checks to stream scope #724 final head b4a311662dbc2250d42a2e3331aa65dbdf898703 的 PR Fast run 35616693360 中,fast job 成功,只有 allocation-gate job 106389543642 失败。
    • 失败 case: rpc-add-sharedmemory-c1。
    • observed median: 1324.768 B/op,低于 allowed median 1450.000 B/op。
    • observed spread: 53.558 B/op,略高于 allowed spread 50.000 B/op,因此 stability check 失败。
    • 同一 gate 的 rpc-add-sharedmemory-c8、rpc-oneway-sharedmemory-c1、send-pump-idle-wake-balanced 均通过;allocation gate harness self-tests 也通过。
    • 同一 head 的 PR Package Smoke 和 CodeQL 成功。

    变更相关性:

    • PR perf: hoist streaming codec capability checks to stream scope #724 修改的是 streaming codec capability hoisting 与对应 regression tests,不涉及 rpc-add-sharedmemory-c1 的 unary/shared-memory allocation path。
    • 当前失败不是 median allocation budget regression,而是单次 sample spread 超阈值 3.558 B/op,初步按 runner/perf-noise stability flake 跟踪。

    证据 / run:

    状态: 已记录;原样重跑 failed job 验证。

  30. SunSi12138 commented on Oct 9, 2026

    @SunSi12138
    OwnerAuthor

    Allocation c1/c8:记录调查结论,维持未解决

    本轮按维护者决定结束调查,暂不修复,也不继续推进门禁拆分或移除 SharedMemory 的方案。生产实现、预算和正式门禁判定保持原样;rpc-add-sharedmemory-c1/c8 仍为未解决,本 issue 保持开放。

    已验证的机制

    • 范围是 Linux x64、.NET 10.0.12;诊断基线为 e91f82118bb4d67997bc76b8a579d51460a28c1b。本次观测发生在 SharedMemory 的通知管道。
    • 预热后,底层管道读取同步完成时,调用边界分配为 0 B;确实挂起时为 144 B。保留的 trace 采到了 PipeStream.ReadAsyncCore 对应的 AsyncStateMachineBox<int, ...>,采样对象大小均为 144 B;独立强制路径实验验证了相应总分配。这是保存异步执行状态的 Task-backed 对象,不是读取的 1 字节缓冲区。默认 async ValueTask 不自动池化。
    • 在一次完整的 c1 观测中,低/高样本差为 21.504 B/op;多出 600 次未完成的 control-pipe read,少一次 SharedMemory 挂起读取,对应调用边界差 21.520 B/op,残差变化 −0.016 B/op。总 control read 只增加 134 次,说明变化主要体现为同步/挂起的完成比例,不能把 600 全部解释成额外通知。
    • 最后的固定混合实验中,Memory/ValueTask 与 byte-array/Task 两个公开重载在 0/25/50/75/100% 挂起比例下,均测得 0/36/72/108/144 B/read;取消、取消后复用及 pending-read disposal 检查均通过。换公开重载没有消除这项开销。
    • perf(runtime): make the suspended-read wrapper allocation-free (#740) #753 处理的是 SharpLink 外层 ReadOwnershipPipeReader 的约 248 B 包装分配;这里的 144 B 位于 .NET 内部,是不同层次。perf(runtime): make the suspended-read wrapper allocation-free (#740) #753 不会直接消除它。这些数值也不能外推到 Windows 或所有 transport。

    证据与边界

    1. 校准后的 trace / 原始门禁对照
    2. 固定路径计数与 SharedMemory ready/pending 对照(完整证据)
    3. 最后一次公开 PipeStream API 校准(完整证据)
    4. .NET 10.0.12 PipeStream 源码;默认 ValueTask builder。

    上述固定诊断序列保留全部样本,来源和结果已独立复核;+512 B/operation 的负向控制仍正确拒绝 c1/c8。诊断中的真实 RPC 门禁没有重现历史 spread 超限,最后一次校准也没有运行 RPC 门禁。路径计数只覆盖同步调用阶段,不能当作完整异步成本模型;历史超限样本没有对应计数,因此尚未完整证明历史失败根因,不能据此判定那些失败无害或已修复。本轮未找到安全、证据充分的最小修复,并不等于偶发分配原则上无法检测。

    另行完成的 TimedBatchShouldExtendBatchForFrameArrivingBeforeDeadline 已由 #794 合入 dev(72806161a0580a047f79348b72783025501daf79);这是独立的 test-only 修复,不解决这里的 allocation c1/c8。

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions