Skip to content

Add PR-level risk assessment score to the review pipeline #4698

Description

@maruiz93

What happens

Fullsend has no composite risk score for pull requests. Several risk signals exist in isolation — protected-path detection in the review post-script (~18 prefixes), finding severity levels, triage risk classification, and RICE effort scoring — but they are evaluated independently by different agents with no unified model. A PR touching CODEOWNERS, CI workflows, and dependency files is treated the same as a docs-only PR in terms of review routing and effort.

What should happen

A PR-level risk assessment should produce a composite score (e.g., 1–5 scale from minimal to critical) before the review agent runs. The score should:

  • Inform review effort: model selection (sonnet vs opus), sub-agent count, and whether human approval is required
  • Be visible as a GitHub label (e.g., risk/low through risk/critical) and optionally as a PR comment with the score breakdown
  • Gate auto-merge eligibility: only low-risk PRs qualify

The assessment should combine three tiers of signals: metadata-derivable signals (cheap, no LLM), git-history-derived signals (LLM-assisted), and linked-issue context signals (LLM-assisted).

Tier 1: Metadata signals (no LLM needed)

Dimension Signal Source
Blast radius Files changed count, lines changed
Path sensitivity Protected path matches, security-sensitive files (.pem, .key, auth/, secrets/)
CI/workflow impact Changes to .github/, Makefiles, Dockerfiles
Dependency risk go.mod/go.sum, package.json, requirements.txt changes
Test coverage Ratio of test files to source files in the diff
Author context Bot vs human, first-time contributor

Tier 2: Git history signals (LLM-assisted)

These dimensions require parsing git history for the files touched by the PR. Raw git data (commit frequency, author counts, commit messages) is collected via CLI, then an LLM interprets the patterns.

Dimension Signal Source
Churn hotspot Files with high recent commit frequency — frequent changes correlate with higher defect rates. LLM contextualizes whether churn is healthy (active development) or concerning (repeated fixes)
Multi-author contention Files recently touched by many different authors. LLM assesses whether the PR author is the primary maintainer or a newcomer to that code area
Recent regression history git log --grep='fix' --grep='revert' on touched files. LLM distinguishes "fix typo" from "fix critical race condition" to weight severity
Change coupling Files that historically change together but aren't all in the PR. LLM assesses whether missing co-changed files are intentional omissions or incomplete changes
Code age / stability Files untouched for months/years carry higher risk when modified (Lindy effect). LLM assesses whether the change is a surgical fix or a broader refactor of stable code
Revert frequency Prior reverts on touched files signal fragile areas
Commit message sentiment History of "workaround", "hack", "temporary", "TODO" in commit messages signals technical debt zones where changes carry extra risk

Tier 3: Linked issue context (LLM-assisted)

When the PR links to an issue, the issue provides additional risk signals that complement the code-level analysis.

Dimension Signal Source
Complexity vs scope mismatch If the linked issue has a high RICE Effort score but the PR is small (or vice versa), the mismatch signals risk — oversimplification, incomplete fix, or scope creep
Issue label context Labels like security, breaking-change, or priority/high on the linked issue directly inform risk beyond what file paths alone suggest
Acceptance criteria coverage LLM compares the issue's requirements/acceptance criteria against what the PR actually changes, flagging potential gaps
Discussion history Extensive debate, unresolved questions, or conflicting opinions on the issue indicate the change is contentious and more likely to need revision
Issue age and staleness Issues open for months may have accumulated context drift — the codebase has evolved since the issue was filed, increasing risk of assumptions that no longer hold

Context

The prioritize agent already implements a RICE scoring model for issues (agents/agents/prioritize.md) with structured JSON output, GitHub Projects field updates, and issue comments. This proposal mirrors that pattern at the PR level.

The review agent's post-review script (agents/scripts/post-review.sh, lines 148–213) already implements protected-path blocking as a post-review gate — the most mature existing risk signal. The pre-review script (pre-review.sh) currently only validates inputs and PR state, making it the natural insertion point for a pre-review risk computation.

Architecture recommendation: Implement in phases:

  1. Pre-review metadata step — Tier 1 signals computed in pre-review.sh with no LLM cost. Produces a baseline risk score and label.
  2. LLM-assisted git history analysis — Tier 2 signals as a lightweight agent or review sub-agent that parses git log output for touched files and refines the risk score.
  3. Linked issue context analysis — Tier 3 signals fetched via gh issue view on the linked issue (if any) and assessed by the same LLM step as Tier 2.

The combined score maps to risk levels:

  • 1–2 (low): Sonnet-only review, auto-merge eligible
  • 3 (medium): Standard review (current behavior)
  • 4 (high): Opus-only review, human approval required
  • 5 (critical): Opus-only review, senior reviewer required, blocks merge

Related issues:

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

agent/reviewReview agentcomponent/dispatchWorkflow dispatch and triggerscomponent/harnessAgent harness, config, and skills loadingfeatureFeature-category issue awaiting human prioritizationpriority/mediumNormal priority, plan for next cycletriagedTriaged but awaiting human prioritizationtype/featureNew capability request

Type

No type

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions