Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
42 changes: 42 additions & 0 deletions .github/ISSUE_TEMPLATE/diagnosis_report.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,42 @@
---
name: Session Diagnosis Report
about: A report produced by the diagnosing-superpowers skill from a real session transcript
labels: bug, automated-issue-report
---

<!--
This template is for reports prepared by the diagnosing-superpowers skill.
The skill fills the sections below from the session transcript and hands
you a prefilled link; review every line before you submit, and attach the
scrubbed bundle if you built one. For anything else, use Bug Report.
-->

- [ ] I searched existing issues and this is not a duplicate

## Environment (required)

| Field | Value |
|-------|-------|
| Superpowers version | |
| Harness (Claude Code, Cursor, etc.) | |
| Harness version | |
| Your model + version | |
| All plugins installed | |
| OS + shell | |

## Is this a Superpowers issue or a platform issue?

- [ ] I confirmed this issue does not occur without Superpowers installed

## What happened?

## Steps to reproduce
1.
2.
3.

## Expected behavior

## Actual behavior

## Debug log or conversation transcript
8 changes: 8 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,6 +22,7 @@ Superpowers is a complete software development methodology for your coding agent
- [Pi](#pi)
- [Hermes Agent](#hermes-agent)
- [The Basic Workflow](#the-basic-workflow)
- [When Something Goes Wrong](#when-something-goes-wrong)
- [Community](#community)
- [What's Inside](#whats-inside)
- [Philosophy](#philosophy)
Expand Down Expand Up @@ -276,6 +277,12 @@ turn loses the bootstrap — start a fresh session if skills stop triggering.

**The agent checks for relevant skills before any task.** Mandatory workflows, not suggestions.

## When Something Goes Wrong

Sometimes a session misbehaves: a skill fires when it shouldn't, stays silent when it should, or the agent ignores its plan, repeats work, or burns more tokens than you'd expect. Ask your coding agent to "figure out what went wrong with superpowers in this session" and it will invoke the **diagnosing-superpowers** skill. To examine an earlier session, name it: "figure out what went wrong with superpowers in session `<id>`".

The skill reads the session transcript, reports what happened with line-level evidence, and, if you want, packages a scrubbed bundle for a bug report.

## Community

Superpowers is built by [Jesse Vincent](https://blog.fsck.com) and the rest of the folks at [Prime Radiant](https://primeradiant.com).
Expand All @@ -294,6 +301,7 @@ Superpowers is built by [Jesse Vincent](https://blog.fsck.com) and the rest of t
**Debugging**
- **systematic-debugging** - 4-phase root cause process (includes root-cause-tracing, defense-in-depth, condition-based-waiting techniques)
- **verification-before-completion** - Ensure it's actually fixed
- **diagnosing-superpowers** - Work out what went wrong in a session, with evidence; export a scrubbed bundle or file an issue

**Collaboration**
- **brainstorming** - Socratic design refinement
Expand Down
1,512 changes: 1,512 additions & 0 deletions docs/superpowers/plans/2026-08-27-diagnosing-superpowers.md

Large diffs are not rendered by default.

530 changes: 530 additions & 0 deletions docs/superpowers/specs/2026-08-27-diagnosing-superpowers-design.md

Large diffs are not rendered by default.

1 change: 1 addition & 0 deletions docs/testing.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,6 +18,7 @@ Live in `tests/`. Currently:
- `tests/claude-code/test-subagent-driven-development-integration.sh` — extended SDD integration with token analysis (drill covers the YAGNI subset; bash adds commit-count, Claude Code task-tracking, and token telemetry assertions).
- `tests/claude-code/test-worktree-native-preference.sh` — RED-GREEN-REFACTOR validation for worktree skill (drill covers the PRESSURE phase; bash also covers RED/GREEN baselines).
- `tests/explicit-skill-requests/` — Haiku-specific, multi-turn, and skill-name-prompted tests not covered by drill.
- `tests/diagnosing-superpowers/test-skill-structure.sh` — structural checks for the diagnosing-superpowers skill (frontmatter, referenced files, leak scan, word budget); behavior-scenario eval records are kept by the maintainer outside the repo.

Run plugin tests via the relevant directory's `run-*.sh` or `npm test`.

Expand Down
118 changes: 118 additions & 0 deletions skills/diagnosing-superpowers/SKILL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,118 @@
---
name: diagnosing-superpowers
description: Use when a superpowers session went wrong and your human partner wants to know why — repeated work, ignored plans, stumbles, poor results, a skill that didn't fire, "it took too long", "why is it so expensive", "what is it doing" — or wants to build a bug report for the superpowers maintainers, for the current session or a past one identified by id or path, on any harness.
---

# Diagnosing Superpowers

## Overview

Pin down with your human partner what went wrong in a session, read the
transcripts on disk, and report what happened with evidence. You report;
you do not diagnose superpowers. Whoever triages the bundle or the issue
decides whether superpowers changes.

**Core principle:** Every finding cites `path:line`. No citation, no
finding. Every number comes from the transcript or from a command you ran,
never from memory.

## Workflow

Create a todo per step. Steps 5–7 run only on their stated condition.

1. **Problem intake.** Ask one question at a time until you can write a
statement naming the session(s), the turn range if known, what your
partner expected, what happened, and the observable they care about
(wall-clock, tokens, repeated actions, one specific action). "It took
too long" is a complaint, not a problem statement. Note whether the
goal is a superpowers bug report.
2. **Locate.** Resolve each session to verified absolute filesystem paths using
`references/session-discovery.md`. Confirm a past session by quoting its
first prompt and timestamp, and list every candidate you rejected with the
reason, or "none". Enumerate subagent transcripts. Create
`~/.superpowers/diagnosing-superpowers/<session-id>/`, tell your
partner the path, and fill `templates/case.md` there, including the
superpowers install root, version, git sha, and a sha1 for every skill
file the session read or had injected.
3. **Triage.** Read the region around the reported problem yourself. Then
dispatch one analyst subagent per dimension in parallel, each given the
case file path, `prompts/analyst-common.md`, and one dimension file from
`prompts/`: `skill-timeline.md`,
`plan-adherence.md`, `repeated-work.md`, `stumbles.md`,
`quality-evidence.md`, `request-conflicts.md`, `cost-and-time.md`.
Split a dimension by turn range when the transcript is long. Discard
any returned finding without `path:line`.
4. **Report.** Fill every section of `templates/report.md` in order, write
it to the workspace, show it, and give the path.
5. **GitHub issues** — when report §7 says possible or likely, or your
partner asks. Search open and closed issues for the symptoms per
`references/github-issues.md`. Show matches and suggest adding the
report to the closest. If none match, fill `templates/issue.md`, write
it to the workspace, show the exact text, and create the issue only
after approval. `gh` cannot attach files; if a bundle exists, give
your partner its path to attach in the browser.
6. **Export** — only when your partner asks for a bundle; never build one
unprompted. If the intake goal was a bug report, say once that a
scrubbed bundle is available on request, then wait. Ask the redaction
level, stating what each includes: skeleton (no tool-result bodies),
evidence (bodies only for cited events), full. Build the bundle per
`templates/bundle-README.md`, dispatch `prompts/scrub.md`, then
`prompts/scrub-audit.md`, repeating both until the audit returns CLEAN.
Show the scrub log and file list; archive (`zip -r` or `tar -czf`)
only after approval. With the archive path, state what it contains,
point at the scrub log for what was replaced, and say scrubbing can
miss things: they must review every file before sharing it.
7. **Similar sessions** — when asked. Turn confirmed findings into a
signature, list candidates by mtime and size, find marker line numbers,
dispatch `prompts/similar-session.md` per candidate in parallel, and
append report §9.

## Quick reference

All seven analysts always run. This table says which region to read
yourself in step 3 and which findings to lead with in the verdict.

| Complaint | Read first, lead with |
|---|---|
| "It took too long" | cost-and-time, stumbles |
| "Why did it do this extra work?" | repeated-work, plan-adherence |
| "Why is it so expensive?" | cost-and-time |
| "What the hell is it doing?" (still running) | skill-timeline; note in-progress in coverage |
| "It ignored the plan" | plan-adherence, compaction lines first |
| "Skill X never fired" | skill-timeline |

## Hard rules

- **Context safety.** One transcript line can be a megabyte. Follow
`references/context-safety.md` on every session file, every time.
- **Read-only.** Never modify, move, or delete a session file.
- **Exact paths to subagents.** A subagent's "current session" is its
own. Pass absolute paths and ids.
- **Human prompts only.** Hook output, system reminders, and tool results
are not your partner's words. In a subagent transcript, "user" is the
parent agent.
- **No superpowers diagnosis.** Report §7 states involvement and stops.
Never name a defect in a skill or propose a change. Your partner
pressing for a fix does not waive this; point at the issue step and
mention that a bundle is available on request. No advice to your
partner either.
- **Approval gates.** No archive before your partner has seen the scrub
log and file list. No issue or comment before they approve the exact
text.
- **Intake before analysis.** Nothing in steps 2–7 starts until your
partner has answered. If they are away, write the questions and stop.
A statement you reconstructed for them is not an answer. An
already-scoped request — one specific event, what is running now, or
the analysis to run — is itself the statement: answer it, then ask.
A whole-session "why" is a complaint.

## Red Flags

| Thought | Reality |
|---------|---------|
| "The problem is obvious, skip intake" | The problem statement scopes everything. Ask. |
| "They're away, so I'll reconstruct the statement" | You cannot reconstruct what they wanted. Write the questions and stop. |
| "I'll sweep everything now and ask at the end" | An unscoped sweep spends their budget on the wrong question. Ask first. |
| "They want a bug report, so I'll build the bundle now" | The bundle is their session data, packaged. Build it only when they ask for it. |
| "Small, targeted edit, no restructuring needed" | Not your call, however small. Report the evidence; the triager decides. |
| "The price per token is well known" | Numbers you did not compute from the transcript are invented. Cite or drop. |
38 changes: 38 additions & 0 deletions skills/diagnosing-superpowers/prompts/analyst-common.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,38 @@
You are an analyst subagent. You read a coding-agent session transcript on
disk and return findings with evidence. You do not fix anything, you do not
modify any file under the session store, and you do not say what
superpowers should change.

Inputs (from your dispatcher):
- CASE: absolute path of the case file. Read it first. It names the session
files, the discovered sources and record meanings to use, and the
context-safety rules you must follow. Use the recorded meanings rather than
repeating discovery or assuming a harness format.
- RANGE (optional): a turn range or line range. If present, analyze only
that range and say so in your Checked line.

Context safety: follow `references/context-safety.md`, named in CASE, on
every file before reading it, and extract fields with the recorded commands or
queries. "The current session" is not a thing you can look at: use only the
paths in CASE.

Human prompts are the records the case file identifies as human-typed. Hook
output, system reminders, and tool results are not human prompts. In a subagent
transcript, "user" is the parent agent.

Return format (nothing else):

```
## <Dimension> findings
- finding: <one sentence, what happened>
evidence: <absolute path>:<line> — "<quote, at most 200 characters>"
turns: <first human turn>–<last human turn>
confidence: high | medium | low
Checked: <what you examined: files, line ranges, commands used>
```

The dispatcher discards any finding without a `path:line`, so do not
write one. If you found nothing, return `- none found` and the Checked
line.
28 changes: 28 additions & 0 deletions skills/diagnosing-superpowers/prompts/cost-and-time.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,28 @@
Read `prompts/analyst-common.md` first; it gives your role, inputs,
context-safety rules, and the return format. This file adds the dimension.

Dimension: Cost and time

Account for where tokens and wall-clock went.

1. Tokens. Use only the usage records and counter meanings established in the
case file. State whether each counter is incremental or cumulative before
calculating totals; difference cumulative observations without turning a
missing observation into zero. Report the five turns with the largest
supported totals and the supported totals per associated session.
2. Wall-clock. Use the evidenced timestamp fields, event boundaries, and units
recorded in the case file. Report the five longest supported turns and any
gap longer than ten minutes between consecutive events (idle, waiting on an
associated session, or waiting on your human partner; say which only when
the records show it).
3. Largest tool results: use the case file's evidenced tool-result records to
report the ten largest results with their tool and turn. Measure records
before extracting bounded content.
4. Compactions: count and locate records whose meaning as compaction events was
established during discovery. Report available before/after counters and
what the session was doing when each fired; mark unsupported fields absent.
5. Associated sessions: count them and report supported usage, duration, and
dispatching turn for each.
6. Report the turns, subagents, tools, or repeats that dominate the
totals, with numbers. Do not speculate about why a
turn was expensive beyond what the transcript shows.
29 changes: 29 additions & 0 deletions skills/diagnosing-superpowers/prompts/plan-adherence.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,29 @@
Read `prompts/analyst-common.md` first; it gives your role, inputs,
context-safety rules, and the return format. This file adds the dimension.

Dimension: Plan adherence

Recover the plan the session agreed to, then map each plan step to what
happened. "Plan" here means any agreed course of action, not git commits.

1. Find the agreed plan: a design or plan agreed in chat (look for the
assistant text preceding a human "yes/ok/go ahead"), a spec or plan file
written during the session (tool calls that write under `docs/`,
`plans/`, `specs/`, or any file the human named), a todo-list record whose
meaning was established in the case file, or any numbered checklist in
assistant text. Quote each plan step with its `path:line`.
2. Mark structural events between the plan and its execution: compaction
events identified during discovery, resumes, aborted turns, and associated
session dispatches. Note their line numbers; plan drift right after one of
these is a distinct finding.
3. For each plan step, find the tool calls and assistant text that
executed it, or establish that none did. Report:
- steps skipped (no execution found; quote the plan step);
- steps executed out of order (line numbers show the order);
- steps silently changed (execution differs from the plan step in a
way the assistant never announced; quote both);
- steps invented (work done that no plan step covers);
- drift immediately after a structural event (cite the event line and
the first divergent action).
4. If there is no recoverable plan, say so as the only finding, with
the lines you checked.
26 changes: 26 additions & 0 deletions skills/diagnosing-superpowers/prompts/quality-evidence.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,26 @@
Read `prompts/analyst-common.md` first; it gives your role, inputs,
context-safety rules, and the return format. This file adds the dimension.

Dimension: Quality evidence

Judge the process against its own claims. This is not a code review; do
not evaluate the code the session produced.

1. Tests: every test run (commands containing `test`, `pytest`, `npm test`,
`cargo test`, `go test`, `bats`, `bash tests/…`, or the project's runner
named in instruction files) with its result line. Report runs that
failed and what the assistant did next.
2. Verification behind claims: find assistant text claiming done, fixed,
passing, verified, works, complete. For each, look backward in the same
turn for a tool result that shows it (a test run, a command output, a
diff). Report claims with no supporting result in that turn.
3. Commits: every `git commit` with its message; compare each message to
the tool calls in the preceding turn(s). Report commits whose message
claims work that no tool call performed, and work performed that was
never committed when the agreed plan said it would be.
4. Review feedback: where a reviewer (human or subagent) raised points,
find the response. Report points acknowledged but not acted on, and
points dismissed without a stated reason.
5. Acceptance criteria: if the case file's problem statement or the
agreed plan states criteria, report each as met / not met /
not checked with the evidence line.
30 changes: 30 additions & 0 deletions skills/diagnosing-superpowers/prompts/repeated-work.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,30 @@
Read `prompts/analyst-common.md` first; it gives your role, inputs,
context-safety rules, and the return format. This file adds the dimension.

Dimension: Repeated work

Find work the session did more than once.

1. Extract every tool call as `(line, turn, tool, key)` where `key` is: the
file path for reads/edits/writes; the command text for shell calls (strip
trailing whitespace; keep the whole command); the `description` plus the
first 80 characters of the prompt for subagent dispatches; the query for
searches.
2. Group by `(tool, key)` and report the groups at or over threshold:

| Category | Threshold | Exempt |
|---|---|---|
| reads, searches | 3 | |
| edits | 2 | |
| shell commands | 2 | status checks and test runs (`git status`, `ls`, `pwd`, test runners) |
| subagent dispatches | 2 with the same description | |
3. For each group, check whether anything changed between repetitions (a
write to that file, a compaction, a human correction). Say which case
it is; a re-read after an edit is not a finding, a re-read after a
compaction is a finding attributed to the compaction, a re-read with
nothing in between is a finding on its own.
4. Look for re-derived decisions: assistant text that reaches a conclusion
already stated earlier in the session (same file, same design choice,
same command to run). Quote both places.
5. One finding per group, with the first and last line numbers and the
count.
20 changes: 20 additions & 0 deletions skills/diagnosing-superpowers/prompts/request-conflicts.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,20 @@
Read `prompts/analyst-common.md` first; it gives your role, inputs,
context-safety rules, and the return format. This file adds the dimension.

Dimension: Request conflicts

1. List every human prompt with line and turn. For each, extract the
instructions it contains (imperatives, constraints, "don't", "always",
"never", "only", scope statements).
2. Report:
- two human instructions that cannot both be followed (quote both, with
lines), and what the assistant did;
- a human instruction that conflicts with an instruction file loaded in
the session (CLAUDE.md, AGENTS.md, GEMINI.md, or the harness's
equivalent; paths are in the case file), quoting both;
- a human instruction to skip, ignore, or override a step, skill, or
rule, and what happened afterwards;
- an instruction the assistant asked to clarify and the answer, when the
answer changed scope.
3. Do not judge whether your human partner was right. Report the conflict
and the assistant's resolution.
Loading