Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

LLM-based Radiology Report Generation & Evaluation Toolkit (with Web Demo)

A toolkit for generating radiology reports from chest X-ray images using Large Language Models (LLMs), and β€” more importantly β€” evaluating those reports along four complementary dimensions, so a single number never tells the whole story.

🎯 Multi-Dimensional Evaluation

Dimension What it measures How
1. NLG quality Surface-level text similarity to the ground truth BLEU-1/2/3/4, ROUGE-L, METEOR, BERTScore β€” via evaluate_nlg.py
2. Clinical accuracy (model-based) Whether the same 14 CheXpert findings are mentioned with the correct polarity CheXbert label extraction β†’ AUC / F1 / Recall / Specificity, via evaluate_chexbert.py
3. Clinical accuracy (LLM-as-Labeler) Same 14-class agreement, but using an LLM as the labeler β€” no chexbert.pth checkpoint needed, fully API-driven evaluate_llm_as_labeler.py (also available in the Web Demo)
4. Radiology-aware semantics Entity / relation overlap and clinical-term-weighted similarity RadGraph F1, RaTEScore (via evaluate_nlg.py)
5. LLM-as-Judge (qualitative) Per-case clinical-quality scores on 4 rubrics (1–10) + explicit missed / hallucinated findings Web Demo β€” see screenshot below

This NLG β¨― Clinical β¨― Semantic β¨― Judge combination lets you tell apart a report that reads well from one that is clinically correct β€” the two often disagree.

The toolkit also ships a Gradio-based Web Demo for interactive single-image report generation and evaluation, with optional Web-Search RAG and an auto-rendered Knowledge Graph linking findings to retrieved evidence.

While we provide out-of-the-box support for Qwen series models (Qwen2.5-VL, Qwen3-VL, Qwen3.5), the evaluation pipeline is model-agnostic β€” any LLM-generated reports in the supported JSON format can be evaluated.

πŸ“Š Jump to Experimental Results ↓

πŸ“ Project Structure

.
β”œβ”€β”€ app.py                         # 🌐 Web demo (Gradio-based UI)
β”œβ”€β”€ qwen_report_generation.py      # Qwen VL inference for report generation
β”œβ”€β”€ evaluate_chexbert.py           # CheXbert-based clinical accuracy evaluation
β”œβ”€β”€ evaluate_nlg.py                # NLG metrics (BLEU, ROUGE, METEOR, BERTScore, etc.)
β”œβ”€β”€ evaluate_llm_as_labeler.py     # LLM-as-labeler clinical accuracy evaluation
β”œβ”€β”€ CheXbert/                      # CheXbert label extraction module
β”‚   └── src/
β”‚       β”œβ”€β”€ label.py
β”‚       β”œβ”€β”€ constants.py
β”‚       β”œβ”€β”€ utils.py
β”‚       β”œβ”€β”€ bert_tokenizer.py
β”‚       β”œβ”€β”€ models/
β”‚       β”‚   └── bert_labeler.py
β”‚       └── datasets_chexbert/
β”‚           └── unlabeled_dataset.py
β”œβ”€β”€ requirements.txt
└── README.md

πŸš€ Installation

pip install -r requirements.txt

# For the web demo
pip install gradio openai

For optional clinical metrics:

# RadGraph F1 (requires model download)
pip install radgraph

# RaTEScore
pip install ratescore

πŸ“‹ Prerequisites

Data Format

Input dataset (test_dataset.json):

{
  "test": [
    {
      "id": "sample_001",
      "role": "user",
      "content": [
        {"type": "image", "image": "/path/to/chest_xray.jpg"},
        {"type": "text", "text": "Please generate a radiology report for this chest X-ray."}
      ]
    }
  ]
}

Annotation file (annotation.json) for evaluation:

{
  "test": [
    {
      "id": "sample_001",
      "report": "No acute cardiopulmonary abnormality. The heart size is normal..."
    }
  ]
}

CheXbert Checkpoint

Download the CheXbert checkpoint from:

Place it at ./checkpoints/chexbert.pth (or specify via --chexbert_checkpoint).

🌐 Web Demo

A simple web interface for single-image report generation and evaluation:

Web Demo β€” Main Interface

pip install gradio openai
python app.py

πŸ€– LLM-as-Judge (clinical-quality scoring)

Reuses your API to ask another LLM to grade the generated report on 4 clinical dimensions (1–10) and to list missed / hallucinated findings against the ground truth.

LLM-as-Judge Result

🏷️ LLM-as-Labeler (14-class CheXpert agreement)

Reuses your API to extract 14 CheXpert binary labels from both the AI report and the ground truth, then computes class-wise agreement, precision, recall and F1 β€” purely API-based, no chexbert.pth required.

LLM-as-Labeler Result

Then open http://localhost:7860 in your browser.

Features:

  • Upload a chest X-ray image and generate a report instantly
  • API mode: Use any OpenAI-compatible API (OpenAI, vLLM, Ollama, Together AI, etc.)
  • Local mode: Load a HuggingFace model locally (requires GPU)
  • Optionally paste a ground truth report to compute metrics (BLEU, ROUGE-L, METEOR)
  • πŸ”Ž Web Search RAG (optional, API mode): let the LLM autonomously retrieve evidence from authoritative medical sites (Radiopaedia, PhysioNet, PubMed/NCBI, Mayo Clinic, RSNA, ...) before writing the report. No local index, no extra files needed β€” everything happens through the API's built-in web_search tool. Provider is auto-routed from the model name:
    • gpt-* β†’ OpenAI web_search tool (with allowed_domains whitelist)
    • glm-* β†’ Zhipu web_search tool
    • qwen-* β†’ DashScope enable_search parameter
    • claude-* β†’ Anthropic web_search_20250305 tool
    • If the chosen API does not support web search, the report is generated without retrieval and a warning is shown.
  • πŸ“‹ Retrieved Sources panel: when web search is enabled, the URLs / titles / snippets returned by the API are extracted and displayed in a sortable table.
  • πŸ•ΈοΈ RAG Knowledge Graph: a Mermaid graph is automatically rendered showing the relationship between detected radiology findings (e.g. cardiomegaly, pleural effusion) and the retrieved web sources that ground them.

Options:

python app.py --port 7860 --share  # Create a public Gradio link
python app.py --server_name 127.0.0.1  # Localhost only

Supported API providers:

Provider API Base URL
OpenAI (leave empty, uses default)
vLLM (local) http://localhost:8000/v1
Ollama http://localhost:11434/v1
Together AI https://api.together.xyz/v1
Any OpenAI-compatible Your endpoint URL

πŸ“– Usage (Batch / Command Line)

1. Report Generation

Generate radiology reports from chest X-ray images using Qwen VL models:

# Using Qwen2.5-VL-7B (default)
python qwen_report_generation.py \
    --model_name Qwen/Qwen2.5-VL-7B-Instruct \
    --question_file ./data/test_dataset.json \
    --output_file ./results/qwen_output.json \
    --max_tokens 512

# Using Qwen3-VL-8B
python qwen_report_generation.py \
    --model_name Qwen/Qwen3-VL-8B-Instruct \
    --model_type qwen3vl \
    --question_file ./data/test_dataset.json \
    --output_file ./results/qwen3vl_output.json \
    --max_tokens 512

# With thinking mode (Qwen3.5)
python qwen_report_generation.py \
    --model_name Qwen/Qwen3.5-27B \
    --question_file ./data/test_dataset.json \
    --output_file ./results/qwen35_output.json \
    --max_tokens 4096 \
    --enable_thinking

# Resume from interrupted run
python qwen_report_generation.py \
    --model_name Qwen/Qwen2.5-VL-7B-Instruct \
    --question_file ./data/test_dataset.json \
    --output_file ./results/qwen_output.json \
    --resume

Key arguments:

Argument Description Default
--model_name HuggingFace model name or local path Qwen/Qwen2.5-VL-7B-Instruct
--model_type auto or qwen3vl auto
--max_tokens Max new tokens to generate 2048
--enable_thinking Enable thinking mode (Qwen3.5) False
--flash_attn Use Flash Attention 2 False
--resume Resume from existing output False

2. CheXbert Evaluation

Evaluate clinical accuracy using CheXbert label extraction:

python evaluate_chexbert.py \
    --llm_output ./results/qwen_output.json \
    --annotation_json ./data/annotation.json \
    --chexbert_checkpoint ./checkpoints/chexbert.pth \
    --output_dir ./results/chexbert_eval

Output:

  • *_eval_summary.xlsx β€” Per-disease AUC, F1, Recall, Specificity + macro average
  • *_per_sample.csv β€” Per-sample binary predictions and agreements

3. NLG Metrics Evaluation

Evaluate with BLEU, ROUGE-L, METEOR, BERTScore, RadGraph F1, RaTEScore:

# Basic NLG metrics (fast, no GPU needed for BLEU/ROUGE/METEOR)
python evaluate_nlg.py \
    --llm_output ./results/qwen_output.json \
    --annotation_json ./data/annotation.json \
    --output_dir ./results/nlg_eval \
    --metrics bleu,rouge,meteor

# All metrics including BERTScore (needs GPU)
python evaluate_nlg.py \
    --llm_output ./results/qwen_output.json \
    --annotation_json ./data/annotation.json \
    --output_dir ./results/nlg_eval \
    --metrics bleu,rouge,meteor,bertscore

# Full evaluation with clinical metrics
python evaluate_nlg.py \
    --llm_output ./results/qwen_output.json \
    --annotation_json ./data/annotation.json \
    --output_dir ./results/nlg_eval \
    --metrics bleu,rouge,meteor,bertscore,radgraph,ratescore

Available metrics:

Metric Package GPU Required
BLEU-1/2/3/4 nltk ❌
ROUGE-L rouge-score ❌
METEOR nltk ❌
BERTScore bert-score βœ…
RadGraph F1 radgraph βœ…
RaTEScore ratescore βœ…

4. LLM-as-Labeler Evaluation

Use another LLM as a label extractor (alternative to CheXbert):

python evaluate_llm_as_labeler.py \
    --llm_output ./results/qwen_output.json \
    --annotation_json ./data/annotation.json \
    --labeler_model Qwen/Qwen2.5-7B-Instruct \
    --output_dir ./results/llm_labeler_eval

Features:

  • Persistent GT cache: ground truth reports are labeled only once per labeler model
  • Resume support: interrupted runs can be continued
  • Both prediction and GT reports are labeled by the same LLM for fair comparison

πŸ“Š Output Format

The report generation script outputs:

{
  "test": [
    {
      "id": "sample_001",
      "output": "The heart size is normal. The lungs are clear..."
    },
    {
      "id": "sample_002",
      "output": "There is mild cardiomegaly. Small bilateral pleural effusions..."
    }
  ]
}

πŸ“ˆ Experimental Results

Below are our evaluation results comparing different models on the MIMIC-CXR test set.

CheXbert as Labeler

Disease Qwen3.5-27B Qwen3-VL-8B
AUCF1RecallSpec AUCF1RecallSpec
Enlarged Cardiomediastinum0.49550.10970.19210.79880.47580.13030.54240.4092
Cardiomegaly0.69890.65510.91930.47840.58270.57030.89240.273
Lung Opacity0.59880.52380.56430.63320.57720.53590.06940.485
Lung Lesion0.50720.04820.03170.98270.50010.03370.02380.9765
Edema0.68420.44710.64850.71980.56240.31150.46530.6595
Consolidation0.51450.06710.04850.98050.51410.07970.11650.9117
Pneumonia0.55290.15050.14000.96590.50740.03540.020.9948
Atelectasis0.55130.28540.21530.88730.50080.01280.00650.995
Pneumothorax0.62800.31710.26530.99070.50810.03390.02040.9958
Pleural Effusion0.71050.61900.61900.80190.55050.27960.19050.9106
Pleural Other0.51420.05560.02940.99910.4958000.9916
Fracture0.49950.00000.00000.99900.499000.9981
Support Devices0.67830.56880.45470.90190.6450.55740.51280.7772
No Finding0.65200.25570.42420.87970.50230.01440.00760.9971
Macro Average0.59180.29310.32520.85850.53010.18540.24770.8125

Qwen3.5 as Labeler

Disease Qwen3.5-27B Qwen3-VL-8B
AUCF1RecallSpec AUCF1RecallSpec
Enlarged Cardiomediastinum0.61080.36220.65250.56910.54870.30880.58750.5099
Cardiomegaly0.70900.64660.92780.49020.59250.55610.87890.3061
Lung Opacity0.62530.59310.69150.55920.59380.58990.77580.4118
Lung Lesion0.52880.10850.06190.99570.49500.00000.00000.9900
Edema0.70310.54640.65770.74860.57730.39860.52070.6338
Consolidation0.52630.09900.09260.96000.50490.07210.11110.8987
Pneumonia0.56890.12000.18000.95790.50120.02250.02000.9824
Atelectasis0.55470.30800.23240.87710.50460.04460.02350.9856
Pneumothorax0.65330.33730.31820.98850.497900.00000.9958
Pleural Effusion0.72170.62760.62930.81410.56350.31970.22980.8972
Pleural Other0.51230.04820.02500.99950.50850.04040.02500.9920
Fracture0.49980.00000.00000.99950.50000.00000.00001.0000
Support Devices0.73650.80420.86400.60900.68350.73760.74510.6219
No Finding0.82130.29180.85710.78550.78230.31620.71430.8503
Macro Average0.62660.34950.44210.81100.56100.24330.33080.7911

Difference (Qwen3.5 as Labeler βˆ’ CheXbert as Labeler)

Disease Qwen3.5-27B Qwen3-VL-8B
Ξ”AUCΞ”F1Ξ”RecallΞ”Spec Ξ”AUCΞ”F1Ξ”RecallΞ”Spec
Enlarged Cardiomediastinum0.11530.25250.4604βˆ’0.22970.07290.17850.04510.1007
Cardiomegaly0.0101βˆ’0.00850.00850.01180.0098βˆ’0.0142βˆ’0.01350.0331
Lung Opacity0.02650.06930.1272βˆ’0.07400.01660.05400.1064βˆ’0.0732
Lung Lesion0.02160.06030.03020.0130βˆ’0.0051βˆ’0.0337βˆ’0.02380.0135
Edema0.01890.09930.00920.02880.01490.08710.0554βˆ’0.0257
Consolidation0.01180.03190.0441βˆ’0.0205βˆ’0.0092βˆ’0.0076βˆ’0.0054βˆ’0.0130
Pneumonia0.0160βˆ’0.03050.0400βˆ’0.0080βˆ’0.0062βˆ’0.01290.0000βˆ’0.0124
Atelectasis0.00340.02260.0171βˆ’0.01020.00380.03180.0170βˆ’0.0094
Pneumothorax0.02530.02020.0529βˆ’0.0022βˆ’0.0102βˆ’0.0339βˆ’0.02040.0000
Pleural Effusion0.01120.00860.01030.01220.01300.04010.0393βˆ’0.0134
Pleural Otherβˆ’0.0019βˆ’0.0074βˆ’0.00440.00040.01270.04040.02500.0004
Fracture0.00030.00000.00000.00050.00100.00000.00000.0019
Support Devices0.05820.23540.4093βˆ’0.29290.03850.18020.2323βˆ’0.1553
No Finding0.16930.03610.4329βˆ’0.09420.28000.30180.7067βˆ’0.1468
Macro Average0.03480.05640.1169βˆ’0.04750.03090.05790.0831βˆ’0.0214

Key Insight: Using Qwen3.5 as the labeler (instead of CheXbert) generally yields higher Recall but slightly lower Specificity. The LLM-based labeler is more sensitive to positive findings mentioned in the generated reports.

πŸ”§ Supported Models

The report generation script provides built-in support for the following Qwen models:

Model --model_name --model_type
Qwen2.5-VL-7B Qwen/Qwen2.5-VL-7B-Instruct auto
Qwen2.5-VL-72B Qwen/Qwen2.5-VL-72B-Instruct auto
Qwen3-VL-8B Qwen/Qwen3-VL-8B-Instruct qwen3vl
Qwen3.5-27B Qwen/Qwen3.5-27B auto

πŸ’‘ Other models: The evaluation scripts (evaluate_chexbert.py, evaluate_nlg.py, evaluate_llm_as_labeler.py) are model-agnostic. As long as your generated reports follow the output JSON format, you can use any VLM/LLM (e.g., GPT-4o, LLaVA-Med, CheXagent, etc.) for report generation and still evaluate with this toolkit.

πŸ“ Notes

  • The CheXbert evaluation uses the U-zeros strategy: only explicit positive (1.0) is treated as positive; everything else (NaN, 0, -1/uncertain) is treated as negative.
  • For the LLM-as-labeler, the same strategy is applied: only explicit positive assertions are labeled as 1.
  • BERTScore uses roberta-large by default (downloaded from HuggingFace). You can specify a local model path with --bertscore_model.
  • The toolkit automatically handles multi-GPU inference via device_map="auto".

⚠️ RadGraph Compatibility

RadGraph depends on an older version of transformers (typically <=4.12.x). It may conflict with the newer transformers version required by Qwen models. It is strongly recommended to create a separate conda/venv environment for RadGraph evaluation:

# Create a separate environment for RadGraph
conda create -n radgraph_env python=3.8 -y
conda activate radgraph_env
pip install radgraph
# Run RadGraph evaluation in this environment
python evaluate_nlg.py \
    --llm_output ./results/qwen_output.json \
    --annotation_json ./data/annotation.json \
    --output_dir ./results/nlg_eval \
    --metrics radgraph

If you encounter errors like ImportError or version conflicts with transformers, this is the expected solution. Other metrics (BLEU, ROUGE, METEOR, BERTScore, RaTEScore) work fine with the latest transformers.

πŸ“„ License

This project uses CheXbert which is subject to its own license. See CheXbert/ for details.

πŸ™ Acknowledgments

  • CheXbert for clinical label extraction
  • Qwen-VL for vision-language models
  • RadGraph for radiology entity extraction

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages