Repository navigation
fix(chat): a finished call whose endpoint reports no usage is billed an estimate, not zero (#1432) - #1438
Conversation
…the estimate, not zero (#1432) Claude-Session: https://claude.ai/code/session_01MkaiVsygNAW7aPgbR7fMEi
Cross-family review (codex), pass 1 |
|
Claude finished @frankbria's task in 10m 10s —— View job GLM precision review — #1438
✅ GLM review: no defects found. Verified against the checked-out repo (not the diff alone):
One non-defect observation (scope of the fix, not a bug in it): the only production constructor of |
|
Reviewed the diff directly (no prior substantive reviews on this PR to reconcile with — just the Correctness
Tests
Scope / docs
No actionable issues found. The known limitations called out in the PR description (estimate over-counts, $0-priced local models still show $0 cost) are accurately described and are inherent to the estimate-based approach rather than bugs in this patch. |
Bot review triage
CI green, no review threads. Merging. |
Closes #1432.
Problem
OpenAIProvider.async_streamasks for usage withstream_options={"include_usage": True}and started its counters at0. Some OpenAI-compatible servers (local ollama/vllm, some proxies) ignore that and never send a usage chunk. The finished call'smessage_stopthen said(0, 0), andStreamingChatAdapter._stream_turnbilled it as nothing. That was less than the same call cut off mid-stream, which gets the #1345 estimate, and the spend was invisible to the per-user daily limit (#1303).Fix
openai.py: the usage counters start atNone, and are set only when a usage chunk arrives. Missing usage reachesmessage_stopas unknown, not0.streaming_chat.py: whenmessage_stopcarries no usage at all (bothNone), the call is billed with the same estimate a cut-off call gets. That estimate moves into one local_estimate(), used by both paths. Reported usage, including a genuine 0, is billed exactly as before.Evidence
tests/api/test_chat_unreported_usage_1432.py: the realOpenAIProviderover a fake SDK stream, through the real producer_run_streaming_adapter.(0, 0)).(123, 45).tests/adaptersplus every test touchingasync_stream/StreamChunk/StreamingChatAdapter/session_chat/message_stop): 438 passed.Cross-family review
codex reviewpass 1: no actionable regressions.Known Limitations
complete()readsresponse.usagedirectly and is unaffected.https://claude.ai/code/session_01MkaiVsygNAW7aPgbR7fMEi