Skip to content

The journal records no sub-agent events, so a delegating run cannot be replayed #6521

Description

@M3gA-Mind

Summary

The durable journal at {workspace}/tinyagents_store/journal carries no sub-agent events at all, and no link from a child run to its parent. A delegating turn therefore cannot be reconstructed from the journal: the delegate's turn span, its iterations, and its tool and model spans are all absent on replay.

Split out of #6419, where the symptom surfaces as a [agent-tracing][journal-shadow] parity divergence warning. PR #6520 makes that warning specific and corrects a doc comment that described the gap wrongly; it does not close this. This is the actual defect.

Evidence

Surveyed the real journal on one developer machine — 871 MB, 6,970 run files.

Zero sub-agent events, across a random sample of 900 run journals. Every event kind that does occur:

count kind
51,985 middleware_started
51,985 middleware_completed
14,834 model_delta
1,006 budget_reserved
895 run_started
755 tools_advertised
748 tools_filtered
592 run_failed
568 middleware_failed
439 model_started
418 model_completed
406 usage_recorded
282 run_completed
203 budget_reconciled
154 tool_started
154 tool_completed
plus a tail of retry_scheduled, handoff_transform_applied, steered, control_applied, unknown_tool_call, limit_reached, invalid_tool_args

Sub-agent events: 0. Journals containing any: 0 of 900.

Meanwhile crates/openhuman-core/src/agent/progress_tracing/journal_projection.rs carries live match arms for AgentEvent::SubAgentStarted, SubAgentCompleted, SubAgentFailed and SubAgentReused. They are dead in production — a consumer waiting on a producer that does not exist.

And the subtree is not reachable indirectly. Every journalled run is its own root:

  • root_run_id == run_id in all 588 run_started records sampled, 0 exceptions.
  • No root_run_id spans more than one run file, across 1,200 files.
  • No run file contains more than one distinct run_id.

So a child run is neither recorded in its parent's journal nor linked from it, and read_run_events(run_id, offset) on the parent cannot reach it by any route.

Why this matters more than the warning

The journal is the only durable record that captures runs when nothing is observing. Langfuse is fail-closed outside staging (LANGFUSE_PUSH_ENVIRONMENTS = &["staging", "development"], and the backend answers 403 FEATURE_DISABLED outside staging), and only the web progress bridge installs a live collector at all — so orchestration, subconscious, cron and meeting turns are untraced live.

That makes the journal the thing you would read to answer any question about a past turn. It is not ground truth for any delegating turn, and delegation is most of what the orchestrator does. Any eval, replay or investigation built on it silently under-reports delegated work — silently being the problem, since the record looks complete.

What would close this

Either is sufficient; the first is better.

  1. Emit sub-agent lifecycle events into the journal — the projection already handles all four kinds, so the spans appear the moment the events do, with no host change.
  2. Populate root_run_id on a child run so a delegating run's tree is discoverable, and have the replay read the tree rather than one run id.

Provider/harness run instrumentation is owned by vendor/tinyagents, so this is an upstream change plus a gitlink bump — filed here per the contribution workflow.

Acceptance criteria

Related

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    agentBuilt-in agents, prompts, orchestration, and agent runtime in src/openhuman/agent/.priority: p1Next. Wrong behaviour a user will hit, or a security weakness behind a condition.

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions