Repository navigation
feat(metrics): populate token usage, add latency percentiles, record … - #25
Open
sandeepyadav1478 wants to merge 1 commit into
Open
sandeepyadav1478 wants to merge 1 commit into
sandeepyadav1478 wants to merge 1 commit into
Conversation
…max_workers The result schema declares GenerationData.prompt_tokens/completion_tokens, but the model is imported by two runners and never constructed, so every published result carries no token data. Both SDKs already report exact usage on each response; llm_client discarded it at the return statement. No new dependency and no local re-tokenization is needed. - llm_client: accumulate SDK-reported usage per eval item. Accounting is held in a ContextVar rather than on the client because one LLMClient is shared by max_workers concurrent items; asyncio copies the context at Task creation, so items cannot bill tokens to each other. Handles both OpenAI (prompt/completion_tokens) and Anthropic (input/output_tokens) field names. - locomo: record answerer tokens per cutoff, separately from whole-item spend. Context-per-query and cost-to-run-the-benchmark are different quantities and folding the judge into the first would misstate it. - metrics: compute_cost_metrics() reports p50/p95 latency, not just a mean. Retrieval latency is right-skewed, and a mean is not comparable to a published p50. It reads search_latency_ms from both locations the benchmarks use (locomo and beam write it at item top level, longmemeval nests it under "retrieval"), so it also works on already-published results. - metadata: record max_workers. Latency was measured under concurrency that the result files never stated, which left every published figure uninterpretable. selfcheck.py covers provider field names, concurrent isolation, untracked calls, percentiles and back-compat with results that predate this change. The repo has no test framework, so it is assert-based and runs on plain python.
sandeepyadav1478
force-pushed
the
feat/token-and-latency-instrumentation
branch
from
August 11, 2026 04:22
01abecf to
6bb44a3
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
…max_workers
The result schema declares GenerationData.prompt_tokens/completion_tokens, but the model is imported by two runners and never constructed, so every published result carries no token data. Both SDKs already report exact usage on each response; llm_client discarded it at the return statement. No new dependency and no local re-tokenization is needed.
selfcheck.py covers provider field names, concurrent isolation, untracked calls, percentiles and back-compat with results that predate this change. The repo has no test framework, so it is assert-based and runs on plain python.