1 of 12

GPT-4V cannot generate radiology reports yet

Yuyang Jiang1, Chacha Chen1, Dang Nyugen1�Benjamin Mervak2, Chenhao Tan1

1 University of Chicago, 2 University of Michigan

2 of 12

3 of 12

4 of 12

RQ1: Direct Report Generation

5 of 12

Why Failure?

Image task

Textual task

6 of 12

Why Failure?

7 of 12

RQ2: GPT-4V cannot interpret medical images meaningfully

8 of 12

9 of 12

RQ3: Report Synthesis Given Groundtruth Conditions

Significant improvements

Gap to ground-truth reports

10 of 12

Additional human evaluation by two radiologists

  1. Human written report are 100% usable, whereas even with groundtruth labels, model generated reports are still not perfect.
  2. Human written reports contains richer and more nuanced information.
  3. Model generated reports have the potential to have better clarity/readability.

Gap to human written reports

11 of 12

Report generation =

Image reasoning

Chest X-rays

<LABEL>

(Cardiomegaly, 0),

(Lung Lesion, 1),

(Lung Opacity, 1),

……

FINDINGS: Hyperinflated with diffuse bilateral opacities. No pleural effusion or pneumothorax. No visible fractures or lytic lesions.

IMPRESSION: Suspected COPD with superimposed infection. No acute disease.

report synthesis

GPT-4V struggles significantly with interpreting chest X-rays meaningfully, which directly impacts its ability to generate reports.

Even when we bypass the bottleneck of image reasoning by providing groundtruth conditions, GPT-4V still underperforms a finetuned LLaMA-2 baseline and fails to match the preferences of radiologists.

Failed terribly

12 of 12

Ongoing: building a radiology foundation model with better data

Key idea: High quality medical data curation for post-training

Find Us at the Poster Session!

Or yuyang2001@uchicago.edu