Papers
arxiv:1705.02315

ChestX-ray8: Hospital-scale Chest X-ray Database and Benchmarks on Weakly-Supervised Classification and Localization of Common Thorax Diseases

Published on May 5, 2017
Authors:
,
,
,
,
,

Abstract

A new chest X-ray database, ChestX-ray8, enables weakly-supervised multi-label image classification and localization for detecting thoracic diseases using deep convolutional neural networks.

The chest X-ray is one of the most commonly accessible radiological examinations for screening and diagnosis of many lung diseases. A tremendous number of X-ray imaging studies accompanied by radiological reports are accumulated and stored in many modern hospitals' Picture Archiving and Communication Systems (PACS). On the other side, it is still an open question how this type of hospital-size knowledge database containing invaluable imaging informatics (i.e., loosely labeled) can be used to facilitate the data-hungry deep learning paradigms in building truly large-scale high precision computer-aided diagnosis (CAD) systems. In this paper, we present a new chest X-ray database, namely "ChestX-ray8", which comprises 108,948 frontal-view X-ray images of 32,717 unique patients with the text-mined eight disease image labels (where each image can have multi-labels), from the associated radiological reports using natural language processing. Importantly, we demonstrate that these commonly occurring thoracic diseases can be detected and even spatially-located via a unified weakly-supervised multi-label image classification and disease localization framework, which is validated using our proposed dataset. Although the initial quantitative results are promising as reported, deep convolutional neural network based "reading chest X-rays" (i.e., recognizing and locating the common disease patterns trained with only image-level labels) remains a strenuous task for fully-automated high precision CAD systems. Data download link: https://nihcc.app.box.com/v/ChestXray-NIHCC

Community

We used ChestX-ray14 as the hardest test in a cross-vendor probing study, and since Table 17 is the baseline we measured against, the numbers belong here.

Averaging the ChestX-ray14 column of Table 17 across the 14 findings gives 0.7451 mean AUROC for the ImageNet-pretrained ResNet-50 on the published split. That is the only figure we compare to. CheXNet (0.8414) and Yao (0.8027) are on a different random partition, so we treat them as not comparable rather than as a target.

On the same official test_list.txt, 25,596 films, patient-disjoint: three frozen multimodal backbones from three vendors, Qwen3-Omni, Gemma-4 and Aria, none fine-tuned, one linear probe each, logits averaged, reach 0.7774. That is +0.0323 over Table 17 with no backbone training and no parameters beyond the three probes.

The margin is not the interesting part. Aria alone scores 0.7080, which is behind Table 17, and pooling it in still raises the average: +0.0124 over the best single backbone, 95% CI [+0.0082, +0.0168] under a patient-clustered bootstrap. Different pretraining corpora appear to leave genuinely different information about the same film.

A probe fitted on one backbone also reads a different backbone's states through a ridge map fitted on training rows only, at a cost of -0.0071. Four of six cross directions beat the probe fitted natively on the target.

The scoping is yours and we inherit it: NLP-mined labels with the precision and recall you document in Table 18, findings visible in the film rather than early detection, linear probes only, three backbones, no clinical validation.

Hidden states and probe weights are published, so the numbers can be re-run rather than taken on trust.

https://huggingface.co/blog/RiverRider/frozen-backbones-read-each-other

Sign up or log in to comment

Get this paper in your agent:

hf papers read 1705.02315
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 2

Datasets citing this paper 4

Spaces citing this paper 9

Browse 9 spaces citing this paper

Collections including this paper 1