GPT-4V cannot generate radiology reports yet
Yuyang Jiang1, Chacha Chen1, Dang Nyugen1�Benjamin Mervak2, Chenhao Tan1
1 University of Chicago, 2 University of Michigan
RQ1: Direct Report Generation
Why Failure?
Image task
Textual task
Why Failure?
RQ2: GPT-4V cannot interpret medical images meaningfully
RQ3: Report Synthesis Given Groundtruth Conditions
Significant improvements
Gap to ground-truth reports
Additional human evaluation by two radiologists
Gap to human written reports
Report generation =
Image reasoning
Chest X-rays
<LABEL>
(Cardiomegaly, 0),
(Lung Lesion, 1),
(Lung Opacity, 1),
……
FINDINGS: Hyperinflated with diffuse bilateral opacities. No pleural effusion or pneumothorax. No visible fractures or lytic lesions.
IMPRESSION: Suspected COPD with superimposed infection. No acute disease.
report synthesis
GPT-4V struggles significantly with interpreting chest X-rays meaningfully, which directly impacts its ability to generate reports.
Even when we bypass the bottleneck of image reasoning by providing groundtruth conditions, GPT-4V still underperforms a finetuned LLaMA-2 baseline and fails to match the preferences of radiologists.
Failed terribly
Ongoing: building a radiology foundation model with better data
Key idea: High quality medical data curation for post-training
Find Us at the Poster Session!
Or yuyang2001@uchicago.edu