A toolkit for generating radiology reports from chest X-ray images using Large Language Models (LLMs), and β more importantly β evaluating those reports along four complementary dimensions, so a single number never tells the whole story.
| Dimension | What it measures | How |
|---|---|---|
| 1. NLG quality | Surface-level text similarity to the ground truth | BLEU-1/2/3/4, ROUGE-L, METEOR, BERTScore β via evaluate_nlg.py |
| 2. Clinical accuracy (model-based) | Whether the same 14 CheXpert findings are mentioned with the correct polarity | CheXbert label extraction β AUC / F1 / Recall / Specificity, via evaluate_chexbert.py |
| 3. Clinical accuracy (LLM-as-Labeler) | Same 14-class agreement, but using an LLM as the labeler β no chexbert.pth checkpoint needed, fully API-driven |
evaluate_llm_as_labeler.py (also available in the Web Demo) |
| 4. Radiology-aware semantics | Entity / relation overlap and clinical-term-weighted similarity | RadGraph F1, RaTEScore (via evaluate_nlg.py) |
| 5. LLM-as-Judge (qualitative) | Per-case clinical-quality scores on 4 rubrics (1β10) + explicit missed / hallucinated findings | Web Demo β see screenshot below |
This NLG β¨― Clinical β¨― Semantic β¨― Judge combination lets you tell apart a report that reads well from one that is clinically correct β the two often disagree.
The toolkit also ships a Gradio-based Web Demo for interactive single-image report generation and evaluation, with optional Web-Search RAG and an auto-rendered Knowledge Graph linking findings to retrieved evidence.
While we provide out-of-the-box support for Qwen series models (Qwen2.5-VL, Qwen3-VL, Qwen3.5), the evaluation pipeline is model-agnostic β any LLM-generated reports in the supported JSON format can be evaluated.
.
βββ app.py # π Web demo (Gradio-based UI)
βββ qwen_report_generation.py # Qwen VL inference for report generation
βββ evaluate_chexbert.py # CheXbert-based clinical accuracy evaluation
βββ evaluate_nlg.py # NLG metrics (BLEU, ROUGE, METEOR, BERTScore, etc.)
βββ evaluate_llm_as_labeler.py # LLM-as-labeler clinical accuracy evaluation
βββ CheXbert/ # CheXbert label extraction module
β βββ src/
β βββ label.py
β βββ constants.py
β βββ utils.py
β βββ bert_tokenizer.py
β βββ models/
β β βββ bert_labeler.py
β βββ datasets_chexbert/
β βββ unlabeled_dataset.py
βββ requirements.txt
βββ README.md
pip install -r requirements.txt
# For the web demo
pip install gradio openaiFor optional clinical metrics:
# RadGraph F1 (requires model download)
pip install radgraph
# RaTEScore
pip install ratescoreInput dataset (test_dataset.json):
{
"test": [
{
"id": "sample_001",
"role": "user",
"content": [
{"type": "image", "image": "/path/to/chest_xray.jpg"},
{"type": "text", "text": "Please generate a radiology report for this chest X-ray."}
]
}
]
}Annotation file (annotation.json) for evaluation:
{
"test": [
{
"id": "sample_001",
"report": "No acute cardiopulmonary abnormality. The heart size is normal..."
}
]
}Download the CheXbert checkpoint from:
Place it at ./checkpoints/chexbert.pth (or specify via --chexbert_checkpoint).
A simple web interface for single-image report generation and evaluation:
pip install gradio openai
python app.pyReuses your API to ask another LLM to grade the generated report on 4 clinical dimensions (1β10) and to list missed / hallucinated findings against the ground truth.
Reuses your API to extract 14 CheXpert binary labels from both the AI report and the ground truth, then computes class-wise agreement, precision, recall and F1 β purely API-based, no chexbert.pth required.
Then open http://localhost:7860 in your browser.
Features:
- Upload a chest X-ray image and generate a report instantly
- API mode: Use any OpenAI-compatible API (OpenAI, vLLM, Ollama, Together AI, etc.)
- Local mode: Load a HuggingFace model locally (requires GPU)
- Optionally paste a ground truth report to compute metrics (BLEU, ROUGE-L, METEOR)
- π Web Search RAG (optional, API mode): let the LLM autonomously retrieve
evidence from authoritative medical sites (Radiopaedia, PhysioNet, PubMed/NCBI,
Mayo Clinic, RSNA, ...) before writing the report. No local index, no extra
files needed β everything happens through the API's built-in
web_searchtool. Provider is auto-routed from the model name:gpt-*β OpenAIweb_searchtool (withallowed_domainswhitelist)glm-*β Zhipuweb_searchtoolqwen-*β DashScopeenable_searchparameterclaude-*β Anthropicweb_search_20250305tool- If the chosen API does not support web search, the report is generated without retrieval and a warning is shown.
- π Retrieved Sources panel: when web search is enabled, the URLs / titles / snippets returned by the API are extracted and displayed in a sortable table.
- πΈοΈ RAG Knowledge Graph: a Mermaid graph is automatically rendered showing the relationship between detected radiology findings (e.g. cardiomegaly, pleural effusion) and the retrieved web sources that ground them.
Options:
python app.py --port 7860 --share # Create a public Gradio link
python app.py --server_name 127.0.0.1 # Localhost onlySupported API providers:
| Provider | API Base URL |
|---|---|
| OpenAI | (leave empty, uses default) |
| vLLM (local) | http://localhost:8000/v1 |
| Ollama | http://localhost:11434/v1 |
| Together AI | https://api.together.xyz/v1 |
| Any OpenAI-compatible | Your endpoint URL |
Generate radiology reports from chest X-ray images using Qwen VL models:
# Using Qwen2.5-VL-7B (default)
python qwen_report_generation.py \
--model_name Qwen/Qwen2.5-VL-7B-Instruct \
--question_file ./data/test_dataset.json \
--output_file ./results/qwen_output.json \
--max_tokens 512
# Using Qwen3-VL-8B
python qwen_report_generation.py \
--model_name Qwen/Qwen3-VL-8B-Instruct \
--model_type qwen3vl \
--question_file ./data/test_dataset.json \
--output_file ./results/qwen3vl_output.json \
--max_tokens 512
# With thinking mode (Qwen3.5)
python qwen_report_generation.py \
--model_name Qwen/Qwen3.5-27B \
--question_file ./data/test_dataset.json \
--output_file ./results/qwen35_output.json \
--max_tokens 4096 \
--enable_thinking
# Resume from interrupted run
python qwen_report_generation.py \
--model_name Qwen/Qwen2.5-VL-7B-Instruct \
--question_file ./data/test_dataset.json \
--output_file ./results/qwen_output.json \
--resumeKey arguments:
| Argument | Description | Default |
|---|---|---|
--model_name |
HuggingFace model name or local path | Qwen/Qwen2.5-VL-7B-Instruct |
--model_type |
auto or qwen3vl |
auto |
--max_tokens |
Max new tokens to generate | 2048 |
--enable_thinking |
Enable thinking mode (Qwen3.5) | False |
--flash_attn |
Use Flash Attention 2 | False |
--resume |
Resume from existing output | False |
Evaluate clinical accuracy using CheXbert label extraction:
python evaluate_chexbert.py \
--llm_output ./results/qwen_output.json \
--annotation_json ./data/annotation.json \
--chexbert_checkpoint ./checkpoints/chexbert.pth \
--output_dir ./results/chexbert_evalOutput:
*_eval_summary.xlsxβ Per-disease AUC, F1, Recall, Specificity + macro average*_per_sample.csvβ Per-sample binary predictions and agreements
Evaluate with BLEU, ROUGE-L, METEOR, BERTScore, RadGraph F1, RaTEScore:
# Basic NLG metrics (fast, no GPU needed for BLEU/ROUGE/METEOR)
python evaluate_nlg.py \
--llm_output ./results/qwen_output.json \
--annotation_json ./data/annotation.json \
--output_dir ./results/nlg_eval \
--metrics bleu,rouge,meteor
# All metrics including BERTScore (needs GPU)
python evaluate_nlg.py \
--llm_output ./results/qwen_output.json \
--annotation_json ./data/annotation.json \
--output_dir ./results/nlg_eval \
--metrics bleu,rouge,meteor,bertscore
# Full evaluation with clinical metrics
python evaluate_nlg.py \
--llm_output ./results/qwen_output.json \
--annotation_json ./data/annotation.json \
--output_dir ./results/nlg_eval \
--metrics bleu,rouge,meteor,bertscore,radgraph,ratescoreAvailable metrics:
| Metric | Package | GPU Required |
|---|---|---|
| BLEU-1/2/3/4 | nltk | β |
| ROUGE-L | rouge-score | β |
| METEOR | nltk | β |
| BERTScore | bert-score | β |
| RadGraph F1 | radgraph | β |
| RaTEScore | ratescore | β |
Use another LLM as a label extractor (alternative to CheXbert):
python evaluate_llm_as_labeler.py \
--llm_output ./results/qwen_output.json \
--annotation_json ./data/annotation.json \
--labeler_model Qwen/Qwen2.5-7B-Instruct \
--output_dir ./results/llm_labeler_evalFeatures:
- Persistent GT cache: ground truth reports are labeled only once per labeler model
- Resume support: interrupted runs can be continued
- Both prediction and GT reports are labeled by the same LLM for fair comparison
The report generation script outputs:
{
"test": [
{
"id": "sample_001",
"output": "The heart size is normal. The lungs are clear..."
},
{
"id": "sample_002",
"output": "There is mild cardiomegaly. Small bilateral pleural effusions..."
}
]
}Below are our evaluation results comparing different models on the MIMIC-CXR test set.
| Disease | Qwen3.5-27B | Qwen3-VL-8B | ||||||
|---|---|---|---|---|---|---|---|---|
| AUC | F1 | Recall | Spec | AUC | F1 | Recall | Spec | |
| Enlarged Cardiomediastinum | 0.4955 | 0.1097 | 0.1921 | 0.7988 | 0.4758 | 0.1303 | 0.5424 | 0.4092 |
| Cardiomegaly | 0.6989 | 0.6551 | 0.9193 | 0.4784 | 0.5827 | 0.5703 | 0.8924 | 0.273 |
| Lung Opacity | 0.5988 | 0.5238 | 0.5643 | 0.6332 | 0.5772 | 0.5359 | 0.0694 | 0.485 |
| Lung Lesion | 0.5072 | 0.0482 | 0.0317 | 0.9827 | 0.5001 | 0.0337 | 0.0238 | 0.9765 |
| Edema | 0.6842 | 0.4471 | 0.6485 | 0.7198 | 0.5624 | 0.3115 | 0.4653 | 0.6595 |
| Consolidation | 0.5145 | 0.0671 | 0.0485 | 0.9805 | 0.5141 | 0.0797 | 0.1165 | 0.9117 |
| Pneumonia | 0.5529 | 0.1505 | 0.1400 | 0.9659 | 0.5074 | 0.0354 | 0.02 | 0.9948 |
| Atelectasis | 0.5513 | 0.2854 | 0.2153 | 0.8873 | 0.5008 | 0.0128 | 0.0065 | 0.995 |
| Pneumothorax | 0.6280 | 0.3171 | 0.2653 | 0.9907 | 0.5081 | 0.0339 | 0.0204 | 0.9958 |
| Pleural Effusion | 0.7105 | 0.6190 | 0.6190 | 0.8019 | 0.5505 | 0.2796 | 0.1905 | 0.9106 |
| Pleural Other | 0.5142 | 0.0556 | 0.0294 | 0.9991 | 0.4958 | 0 | 0 | 0.9916 |
| Fracture | 0.4995 | 0.0000 | 0.0000 | 0.9990 | 0.499 | 0 | 0 | 0.9981 |
| Support Devices | 0.6783 | 0.5688 | 0.4547 | 0.9019 | 0.645 | 0.5574 | 0.5128 | 0.7772 |
| No Finding | 0.6520 | 0.2557 | 0.4242 | 0.8797 | 0.5023 | 0.0144 | 0.0076 | 0.9971 |
| Macro Average | 0.5918 | 0.2931 | 0.3252 | 0.8585 | 0.5301 | 0.1854 | 0.2477 | 0.8125 |
| Disease | Qwen3.5-27B | Qwen3-VL-8B | ||||||
|---|---|---|---|---|---|---|---|---|
| AUC | F1 | Recall | Spec | AUC | F1 | Recall | Spec | |
| Enlarged Cardiomediastinum | 0.6108 | 0.3622 | 0.6525 | 0.5691 | 0.5487 | 0.3088 | 0.5875 | 0.5099 |
| Cardiomegaly | 0.7090 | 0.6466 | 0.9278 | 0.4902 | 0.5925 | 0.5561 | 0.8789 | 0.3061 |
| Lung Opacity | 0.6253 | 0.5931 | 0.6915 | 0.5592 | 0.5938 | 0.5899 | 0.7758 | 0.4118 |
| Lung Lesion | 0.5288 | 0.1085 | 0.0619 | 0.9957 | 0.4950 | 0.0000 | 0.0000 | 0.9900 |
| Edema | 0.7031 | 0.5464 | 0.6577 | 0.7486 | 0.5773 | 0.3986 | 0.5207 | 0.6338 |
| Consolidation | 0.5263 | 0.0990 | 0.0926 | 0.9600 | 0.5049 | 0.0721 | 0.1111 | 0.8987 |
| Pneumonia | 0.5689 | 0.1200 | 0.1800 | 0.9579 | 0.5012 | 0.0225 | 0.0200 | 0.9824 |
| Atelectasis | 0.5547 | 0.3080 | 0.2324 | 0.8771 | 0.5046 | 0.0446 | 0.0235 | 0.9856 |
| Pneumothorax | 0.6533 | 0.3373 | 0.3182 | 0.9885 | 0.4979 | 0 | 0.0000 | 0.9958 |
| Pleural Effusion | 0.7217 | 0.6276 | 0.6293 | 0.8141 | 0.5635 | 0.3197 | 0.2298 | 0.8972 |
| Pleural Other | 0.5123 | 0.0482 | 0.0250 | 0.9995 | 0.5085 | 0.0404 | 0.0250 | 0.9920 |
| Fracture | 0.4998 | 0.0000 | 0.0000 | 0.9995 | 0.5000 | 0.0000 | 0.0000 | 1.0000 |
| Support Devices | 0.7365 | 0.8042 | 0.8640 | 0.6090 | 0.6835 | 0.7376 | 0.7451 | 0.6219 |
| No Finding | 0.8213 | 0.2918 | 0.8571 | 0.7855 | 0.7823 | 0.3162 | 0.7143 | 0.8503 |
| Macro Average | 0.6266 | 0.3495 | 0.4421 | 0.8110 | 0.5610 | 0.2433 | 0.3308 | 0.7911 |
| Disease | Qwen3.5-27B | Qwen3-VL-8B | ||||||
|---|---|---|---|---|---|---|---|---|
| ΞAUC | ΞF1 | ΞRecall | ΞSpec | ΞAUC | ΞF1 | ΞRecall | ΞSpec | |
| Enlarged Cardiomediastinum | 0.1153 | 0.2525 | 0.4604 | β0.2297 | 0.0729 | 0.1785 | 0.0451 | 0.1007 |
| Cardiomegaly | 0.0101 | β0.0085 | 0.0085 | 0.0118 | 0.0098 | β0.0142 | β0.0135 | 0.0331 |
| Lung Opacity | 0.0265 | 0.0693 | 0.1272 | β0.0740 | 0.0166 | 0.0540 | 0.1064 | β0.0732 |
| Lung Lesion | 0.0216 | 0.0603 | 0.0302 | 0.0130 | β0.0051 | β0.0337 | β0.0238 | 0.0135 |
| Edema | 0.0189 | 0.0993 | 0.0092 | 0.0288 | 0.0149 | 0.0871 | 0.0554 | β0.0257 |
| Consolidation | 0.0118 | 0.0319 | 0.0441 | β0.0205 | β0.0092 | β0.0076 | β0.0054 | β0.0130 |
| Pneumonia | 0.0160 | β0.0305 | 0.0400 | β0.0080 | β0.0062 | β0.0129 | 0.0000 | β0.0124 |
| Atelectasis | 0.0034 | 0.0226 | 0.0171 | β0.0102 | 0.0038 | 0.0318 | 0.0170 | β0.0094 |
| Pneumothorax | 0.0253 | 0.0202 | 0.0529 | β0.0022 | β0.0102 | β0.0339 | β0.0204 | 0.0000 |
| Pleural Effusion | 0.0112 | 0.0086 | 0.0103 | 0.0122 | 0.0130 | 0.0401 | 0.0393 | β0.0134 |
| Pleural Other | β0.0019 | β0.0074 | β0.0044 | 0.0004 | 0.0127 | 0.0404 | 0.0250 | 0.0004 |
| Fracture | 0.0003 | 0.0000 | 0.0000 | 0.0005 | 0.0010 | 0.0000 | 0.0000 | 0.0019 |
| Support Devices | 0.0582 | 0.2354 | 0.4093 | β0.2929 | 0.0385 | 0.1802 | 0.2323 | β0.1553 |
| No Finding | 0.1693 | 0.0361 | 0.4329 | β0.0942 | 0.2800 | 0.3018 | 0.7067 | β0.1468 |
| Macro Average | 0.0348 | 0.0564 | 0.1169 | β0.0475 | 0.0309 | 0.0579 | 0.0831 | β0.0214 |
Key Insight: Using Qwen3.5 as the labeler (instead of CheXbert) generally yields higher Recall but slightly lower Specificity. The LLM-based labeler is more sensitive to positive findings mentioned in the generated reports.
The report generation script provides built-in support for the following Qwen models:
| Model | --model_name |
--model_type |
|---|---|---|
| Qwen2.5-VL-7B | Qwen/Qwen2.5-VL-7B-Instruct |
auto |
| Qwen2.5-VL-72B | Qwen/Qwen2.5-VL-72B-Instruct |
auto |
| Qwen3-VL-8B | Qwen/Qwen3-VL-8B-Instruct |
qwen3vl |
| Qwen3.5-27B | Qwen/Qwen3.5-27B |
auto |
π‘ Other models: The evaluation scripts (
evaluate_chexbert.py,evaluate_nlg.py,evaluate_llm_as_labeler.py) are model-agnostic. As long as your generated reports follow the output JSON format, you can use any VLM/LLM (e.g., GPT-4o, LLaVA-Med, CheXagent, etc.) for report generation and still evaluate with this toolkit.
- The CheXbert evaluation uses the U-zeros strategy: only explicit positive (1.0) is treated as positive; everything else (NaN, 0, -1/uncertain) is treated as negative.
- For the LLM-as-labeler, the same strategy is applied: only explicit positive assertions are labeled as 1.
- BERTScore uses
roberta-largeby default (downloaded from HuggingFace). You can specify a local model path with--bertscore_model. - The toolkit automatically handles multi-GPU inference via
device_map="auto".
RadGraph depends on an older version of transformers (typically <=4.12.x). It may conflict with the newer transformers version required by Qwen models. It is strongly recommended to create a separate conda/venv environment for RadGraph evaluation:
# Create a separate environment for RadGraph
conda create -n radgraph_env python=3.8 -y
conda activate radgraph_env
pip install radgraph
# Run RadGraph evaluation in this environment
python evaluate_nlg.py \
--llm_output ./results/qwen_output.json \
--annotation_json ./data/annotation.json \
--output_dir ./results/nlg_eval \
--metrics radgraphIf you encounter errors like ImportError or version conflicts with transformers, this is the expected solution. Other metrics (BLEU, ROUGE, METEOR, BERTScore, RaTEScore) work fine with the latest transformers.
This project uses CheXbert which is subject to its own license. See CheXbert/ for details.


