Skip to content

feat(metrics): populate token usage, add latency percentiles, record … - #25

Open
sandeepyadav1478 wants to merge 1 commit into
mem0ai:mainfrom
sandeepyadav1478:feat/token-and-latency-instrumentation
Open

sandeepyadav1478 wants to merge 1 commit into
mem0ai:mainfrom
sandeepyadav1478:feat/token-and-latency-instrumentation

Conversation

@sandeepyadav1478

Copy link
Copy Markdown

…max_workers

The result schema declares GenerationData.prompt_tokens/completion_tokens, but the model is imported by two runners and never constructed, so every published result carries no token data. Both SDKs already report exact usage on each response; llm_client discarded it at the return statement. No new dependency and no local re-tokenization is needed.

  • llm_client: accumulate SDK-reported usage per eval item. Accounting is held in a ContextVar rather than on the client because one LLMClient is shared by max_workers concurrent items; asyncio copies the context at Task creation, so items cannot bill tokens to each other. Handles both OpenAI (prompt/completion_tokens) and Anthropic (input/output_tokens) field names.
  • locomo: record answerer tokens per cutoff, separately from whole-item spend. Context-per-query and cost-to-run-the-benchmark are different quantities and folding the judge into the first would misstate it.
  • metrics: compute_cost_metrics() reports p50/p95 latency, not just a mean. Retrieval latency is right-skewed, and a mean is not comparable to a published p50. It reads search_latency_ms from both locations the benchmarks use (locomo and beam write it at item top level, longmemeval nests it under "retrieval"), so it also works on already-published results.
  • metadata: record max_workers. Latency was measured under concurrency that the result files never stated, which left every published figure uninterpretable.

selfcheck.py covers provider field names, concurrent isolation, untracked calls, percentiles and back-compat with results that predate this change. The repo has no test framework, so it is assert-based and runs on plain python.

…max_workers

The result schema declares GenerationData.prompt_tokens/completion_tokens, but
the model is imported by two runners and never constructed, so every published
result carries no token data. Both SDKs already report exact usage on each
response; llm_client discarded it at the return statement. No new dependency and
no local re-tokenization is needed.

- llm_client: accumulate SDK-reported usage per eval item. Accounting is held in
  a ContextVar rather than on the client because one LLMClient is shared by
  max_workers concurrent items; asyncio copies the context at Task creation, so
  items cannot bill tokens to each other. Handles both OpenAI
  (prompt/completion_tokens) and Anthropic (input/output_tokens) field names.
- locomo: record answerer tokens per cutoff, separately from whole-item spend.
  Context-per-query and cost-to-run-the-benchmark are different quantities and
  folding the judge into the first would misstate it.
- metrics: compute_cost_metrics() reports p50/p95 latency, not just a mean.
  Retrieval latency is right-skewed, and a mean is not comparable to a published
  p50. It reads search_latency_ms from both locations the benchmarks use (locomo
  and beam write it at item top level, longmemeval nests it under "retrieval"),
  so it also works on already-published results.
- metadata: record max_workers. Latency was measured under concurrency that the
  result files never stated, which left every published figure uninterpretable.

selfcheck.py covers provider field names, concurrent isolation, untracked calls,
percentiles and back-compat with results that predate this change. The repo has
no test framework, so it is assert-based and runs on plain python.
@sandeepyadav1478
sandeepyadav1478 force-pushed the feat/token-and-latency-instrumentation branch from 01abecf to 6bb44a3 Compare August 11, 2026 04:22
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant