Multicentre performance and consistency of two deep learning models for malignancy probability estimation of incidental pulmonary nodules

authors
Dinnessen et al.
year
2026
category
biomedical-imaging
pdf
PDF
source
Source
업데이트
2026-08-23

이 논문의 관계도

List view

요약

이 연구는 incidental pulmonary nodules에 대해 malignancy probability estimation을 수행하는 두 가지 deep learning (DL) 모델(DL1: screening data only, DL2: screening + clinical data)의 multicentre 성능과 일관성을 평가했다. retrospective case-control design으로 3개 Dutch centre에서 수집된 269개의 nodule(89 malignant, 180 benign)을 대상으로 분석했으며, nodule size bucket(5–10 mm, 10–15 mm, 15–30 mm)별로 stratified sampling을 적용했다. DL1과 DL2는 pooled dataset에서 AUC가 각각 0.74와 0.72로, 기존 clinical standard인 Brock model(AUC 0.63)보다 통계적으로 유의하게 높은 discrimination 능력을 보였다(p < 0.01). 고정된 sensitivity(77.5%) 하에서 DL 모델들의 specificity는 60%로 Brock model의 44.4%를 상회했다. centre-stratified 분석 결과, DL 모델들은 scanner manufacturer와 acquisition protocol이 다른 centre 간에서도 일관된 성능을 유지했으며, 특히 non-university teaching hospital인 centre 2와 university hospital인 centre 3에서 Brock model 대비 우월성이 뚜렷했다. 흥미롭게도 clinical data를 추가 학습한 DL2는 screening data만 학습한 DL1보다 성능 향상이 관찰되지 않았으며, 이는 DL1이 이미 sufficient features를 학습했거나 추가 malignant sample(n=127)의 규모가 부족했기 때문으로 해석된다.

방법

Study design과 cohort 구성: 이 연구는 retrospective, multicentre, case-control design을 채택했다. Netherlands 내 3개 centre(2개 university hospital, 1개 non-university teaching hospital)에서 incidental pulmonary nodules을 수집했다. inclusion criteria는 age ≥ 18세, CT scan에서 pulmonary nodule 진단, 그리고 최소 2년 follow-up 후 stability/resolution 확인 또는 definitive diagnostic work-up 완료 환자였다. exclusion criteria로는 prior cancer history, MEN1 syndrome, benign nodular diseases (rheumatoid arthritis, common variable immunodeficiency, granulomatosis with polyangiitis), >10개의 nodule, advanced fibrosis, carcinoid cancer 등이 포함되었다. 또한 tree-in-bud configuration, calcified nodules, typical perifissural nodules, hamartomas containing macroscopic fatty areas, cystic nodules은 제외했다. CT scan eligibility는 full thorax coverage, reconstruction matrix 512×512 또는 1024×1024, slice thickness ≤ 3 mm를 요구했다.

Sampling strategy와 data collection: Nodule size가 lung cancer risk와 강하게 연관되어 있으므로, size-based buckets(5–10 mm, 10–15 mm, 15–30 mm)을 정의하고 각 bucket당 centre별로 10개 malignant nodule과 20개 benign nodule을 목표로 sampling했다. 최종적으로 269개의 nodule(89 malignant, 180 benign)이 분석에 포함되었다. Malignant cases는 pathology records에서 확인했고, benign controls는 Natural Language Processing algorithm을 이용해 radiology reports에서 pulmonary nodule 언급을 검색하거나 PACS keyword search를 통해 식별한 후 manual verification을 거쳤다. Nodule size는 semi-automated software를 사용하여 모든 centre에서 standardized manner로 maximum diameter를 측정했다. Patient data(age, gender)는 Electronic Medical Record System에서 추출했고, lobe location은 radiology report에서 추출하여 local radiologist가 verified했다. Emphysema는 inter-reader variability가 높고 Brock model에서의 contribution이 cohort에 따라 달라지므로 annotation하지 않았다.

Model architecture와 comparison: 세 가지 모델을 비교했다: (1) Brock model (logistic regression based on patient/nodule characteristics), (2) DL1 (screening-trained deep learning model, input: 5×5×5 cm block around nodule, trained on National Lung Cancer Screening Trial data [n=16,077 nodules, 1,249 malignant]), (3) DL2 (updated DL model, trained on screening data + clinical cohort data [127 malignant incidental nodules], benign nodules automatically detected from patients without cancer diagnosis). DL1은 previously validated된 모델이며, DL2는 본 연구에서 처음 validation되었다.

Statistical analysis: Primary endpoint는 malignancy probability estimation의 discrimination performance로, Receiver Operating Characteristic (ROC) curve와 Area Under the Curve (AUC)를 사용하여 평가했다. AUC 비교는 DeLong method를 사용했고, 95% confidence intervals는 bootstrapping (10,000 samples)으로 계산했다. Secondary endpoint로는 fixed sensitivity 하에서의 specificity를 비교했으며, target sensitivity는 Brock model이 British Thoracic Society guidelines에서 사용하는 10% threshold에서의 sensitivity(77.5%)와 동일하게 설정했다. Consistency of discrimination은 centre-stratified ROC analysis로 평가했다. Patient-level sensitivity analysis도 수행하여, patient당 highest risk score nodule만 선택하여 재분석했다. Statistical significance는 p < 0.05로 정의했으며, R (version 4.4.0)을 사용하여 분석을 수행했다.

주요 결과

Primary endpoint (AUC performance): Pooled multicentre dataset에서 DL1의 AUC는 0.74 (95% CI: 0.68–0.80), DL2는 0.72 (95% CI: 0.66–0.79)였으며, Brock model은 0.63 (95% CI: 0.56–0.70)이었다. 두 DL 모델 모두 Brock model보다 통계적으로 유의하게 높은 AUC를 보였다 (both p < 0.01). Patient-level analysis (selecting highest risk score per patient, n=85 malignant, 146 benign)에서도 DL1 (AUC 0.71, 95% CI: 0.65–0.78)과 DL2 (AUC 0.70, 95% CI: 0.63–0.77)는 Brock model (AUC 0.61, 95% CI: 0.54–0.68)보다 우월했다.

Secondary endpoint (Specificity at fixed sensitivity): Brock model의 10% threshold에서 달성된 sensitivity (77.5%)를 target으로 할 때, DL1과 DL2는 각각 11%와 7.8% threshold에서 동일한 sensitivity (77.5%)를 달성했으며, 이때 specificity는 두 DL 모델 모두 60%였다. 반면 Brock model의 specificity는 44.4%에 불과했다. 이는 DL 모델이 동일한 cancer detection rate를 유지하면서 false positive rates를 현저히 낮출 수 있음을 시사한다.

Centre-stratified results:

Nodule size correlation: Brock model은 nodule diameter와 매우 강한 상관관계를 보였다 (Centre 1: r=0.89, Centre 2: r=0.86, Centre 3: r=0.91; all p<0.01). 반면 DL 모델들은 Centre 1에서 strong correlation (DL1/DL2: r=0.73)을 보였으나, Centre 2와 3에서는 weak to moderate correlation (Centre 2: DL1 r=0.40, DL2 r=0.52; Centre 3: DL1 r=0.40, DL2 r=0.59; all p<0.01)을 보였다. 이는 DL 모델이 size 외에도 다른 imaging features를 활용하여 discrimination을 수행함을 시사한다.

DL1 vs DL2 performance: Clinical data를 추가 학습한 DL2는 screening data만 학습한 DL1보다 성능에서 우위를 보이지 않았다. Pooled dataset 및 각 centre별 분석 모두에서 두 모델의 AUC와 specificity는 유사했다. 이는 DL1이 이미 sufficient features를 학습했거나, 추가된 clinical malignant nodules (n=127)의 규모가 screening data (n=1,249)에 비해 상대적으로 작아 meaningful improvement를 제공하지 못했기 때문으로 해석된다.

강점

첫째, multicentre design을 통해 scanner manufacturer (Siemens, Canon, GE, Philips, Toshiba), acquisition protocol (standard, HR, low dose, CTA), slice thickness, matrix size 등의 heterogeneity가 있는 실제 clinical setting에서의 model robustness를 평가했다. 이는 single-centre study나 homogeneous screening cohort 기반 연구의 한계를 극복한다. 둘째, size-stratified case-control sampling을 통해 nodule size라는 강력한 confounder의 영향을 통제함으로써, 모델이 size 외의 imaging features를 얼마나 잘 학습했는지를 더 정확하게 평가할 수 있었다. 셋째, widely used clinical standard인 Brock model과 직접 비교하여 DL 모델의 relative performance를 명확히 제시했다. 넷째, centre-stratified analysis를 통해 모델 성능이 특정 centre나 scanner에 의존하지 않고 일관되게 유지되는지 확인함으로써 generalizability에 대한 근거를 강화했다.

한계

첫째, cancer-enriched dataset (malignant:benign ratio ~1:2)을 사용했으므로, real-world prevalence (much lower malignancy rate)에서의 performance, 특히 positive predictive value와 calibration은 overestimated될 수 있다. 저자들도 prospective validation in representative cohorts with real-world prevalences가 필요함을 명시했다. 둘째, sample size가 relatively small (n=269 nodules)하여, subgroup analysis나 rare nodule types에 대한 generalizability이 제한적이다. 셋째, DL2 모델의 clinical training data 중 malignant nodules이 127개에 불과해, screening data (1,249 malignant)에 비해 insufficient amount로 판단되어 performance improvement가 관찰되지 않았을 가능성이 있다. 넷째, model input이 nodule 주변 5×5×5 cm block으로 제한되어, lung parenchyma의 다른 부분 (예: emphysema, fibrosis elsewhere)에서 얻을 수 있는 contextual information을 활용하지 못했다는 점이 한계로 지적되었다. 다섯째, retrospective design과 case-control sampling strategy로 인해 selection bias가 존재할 수 있으며, 특히 benign controls의 identification method (NLP/keyword search + manual check)에 따라 misclassification 가능성이 배제되지 않는다.

해석

이 연구는 screening-trained DL model (DL1)이 incidental pulmonary nodules에서도 robust하게 작동하며, clinical data를 추가 학습한 모델 (DL2)이 반드시 성능을 향상시키지 않음을 시사한다. 이는 DL model이 이미 sufficient imaging features를 학습했거나, clinical setting에서의 nodule imaging characteristics가 screening setting과 유사할 수 있음을 의미한다. 특히, Brock model이 nodule size에 과도하게 의존하는 반면, DL 모델들은 size 외의 features를 활용하여 discrimination을 수행하므로, size-stratified dataset에서도 우월한 성능을 보였다. 이는 DL model이 radiological features beyond simple diameter (e.g., texture, shape, margin)를 효과적으로 학습했음을 시사한다. LLM Wiki의 oncology/imaging 문헌들과 연결해 보면, 이 연구는 AI-based nodule assessment가 single-centre validation을 넘어 multicentre clinical setting에서도 feasible함을 보여주는 중요한 근거이다. 그러나 real-world prevalence에서의 calibration과 prospective validation이 여전히 필요하며, 특히 low-prevalence setting에서의 false positive rate 관리가 임상 적용의 핵심 과제로 남는다.