Artificial Intelligence-Powered Spatial Analysis of Tumor-Infiltrating Lymphocytes as Complementary Biomarker for Immune Checkpoint Inhibition in Non-Small-Cell Lung Cancer
이 논문의 관계도
List view
Shared Tags이 논문의 tags: AI, ICI, NSCLC, lung cancer, real-world
- Hao 2022Immune checkpoint inhibitor-related pneumonitis in non-small cell lung cancer: A review공유 tag: AI, ICI, NSCLC, lung cancer
- He 2020Predicting response to immunotherapy in advanced non-small-cell lung cancer using tumor mutatio…공유 tag: ICI, NSCLC, lung cancer
- Goldstraw 2016The IASLC Lung Cancer Staging Project: Proposals for Revision of the TNM Stage Groupings in the…공유 tag: AI, NSCLC, lung cancer
- Rocha 2018CD103+CD8+ Lymphocytes Characterize the Immune Infiltration in a Case With Pseudoprogression in…공유 tag: ICI, NSCLC, lung cancer
- Jung 2020Real world data of durvalumab consolidation after chemoradiotherapy in stage III non-small-cell…공유 tag: AI, NSCLC, lung cancer
- +98 more
요약
Park et al.은 advanced NSCLC 환자에서 ICI 치료 반응 및 생존을 예측하기 위해 AI 기반 WSI 분석 도구인 Lunit SCOPE IO를 개발하고 검증하였다. 이 도구는 H&E 염색된 조직 슬라이드에서 cancer epithelium(CE), cancer stroma(CS) 및 tumor-infiltrating lymphocytes(TIL)를 자동 분할 및 정량화하여 세 가지 immune phenotype(IP: inflamed, immune-excluded, immune-desert)을 정의한다. 연구 결과, inflamed IP는 높은 tumor response rate와 연장된 progression-free survival(PFS) 및 overall survival(OS)과 유의미하게 연관되었으며, 기존 pathologist가 판독한 PD-L1 TPS와도 강한 상관관계를 보였다. 이는 AI 기반 spatial analysis가 PD-L1 TPS를 보완하는 objective biomarker로서 임상적 유용성을 가짐을 시사한다.
방법
Lunit SCOPE IO 모델은 25가지 cancer types의 3,166개 H&E-stained WSI(약 $2.83 \times 10^9$ mm² 영역 및 $6.0 \times 10^5$ TIL)를 board-certified pathologist annotation으로 학습하여 CE, CS segmentation 및 TIL detection을 수행하도록 개발되었다. 모델 성능 검증은 internal validation과 external validation(TCGA LUAD/LUSC n=110, ground-truth는 3명의 pathologist consensus)을 통해 이루어졌다. Clinical validation cohort로는 SMC(NSCLC primary tumor n=1,205, ICI treatment outcome n=299)와 SNUBH(NSCLC n=261, ICI treatment outcome n=219)의 두 independent cohort를 사용하였다. WSI는 $1 \text{ mm}^2$ grid로 분할된 후, CE 내 TIL density threshold ($> 106/\text{mm}^2$)와 CS 내 TIL density threshold ($> 357/\text{mm}^2$)를 기준으로 IP를 분류하였다. PD-L1 expression은 Dako PD-L1 IHC 22C3 pharmDx kit를 이용해 TPS로 측정되었으며, clinical outcome(best overall response, PFS, OS)은 RECIST v1.1 기준에 따라 retrospectively 평가되었다. 또한 TCGA genomic data(n=923) 및 multiplex IHC 분석을 통해 IP의 biological validity를 검증하였다.
주요 결과
AI model은 internal validation에서 CE(AUROC 0.9715), CS(AUROC 0.9503), TIL(AUROC 0.9252) segmentation 및 detection 성능을 보였으며, external validation(TCGA n=110)에서도 CE(0.9539), CS(0.9871), TIL(0.9591)로 높은 성능을 유지하였다. 전체 cohort에서 inflamed IP는 44.0%, immune-excluded IP는 37.1%, immune-desert IP는 18.9%로 분포되었다. PD-L1 TPS <1%, 1%-49%, ≥50% 그룹에서 inflamed IP incidence는 각각 31.7%, 42.5%, 56.8%였으며, AI-derived TPS와 pathologist TPS 간 유의한 positive correlation(P < .001)이 확인되었다. Clinical survival analysis 결과, inflamed IP cohort의 median PFS 및 OS는 각각 4.1개월, 24.8개월로 immune-excluded IP(2.2개월, 14.0개월) 및 immune-desert IP(2.4개월, 10.6개월)보다 유의하게 길었다. Genomic analysis에서 inflamed IP는 local immune cytolytic activity(GZMA, PRF1 expression 증가) 및 interferon gamma response 관련 gene signature와 밀접한 연관성을 보였다.
통계 분석
분석 설계 — 이 연구는 AI 기반의 공간 분석을 통해 분류된 immune phenotype(IP)이 advanced NSCLC 환자의 ICI 치료 반응 및 생존에 어떤 예측 가치를 가지는지 확인하기 위해 수행되었다. 연구 설계는 두 가지 독립적인 cohort(SMC cohort, n=299; SNUBH cohort, n=219; 총 N=518)를 활용한 retrospective observational study 형태이다. primary endpoint는 progression-free survival(PFS)과 overall survival(OS)이며, secondary endpoint로는 tumor response rate(TRR) 및 PD-L1 TPS와의 상관관계 분석이 포함되었다. 추적 기간은 명시되지 않았으나, median PFS와 OS가 보고됨으로써 충분한 follow-up이 이루어진 것으로 추정된다.
무엇을 위해 어떤 분석을 썼는가 — AI 모델의 성능(CE, CS segmentation 및 TIL detection)을 평가하기 위해 receiver operating characteristic curves와 area under the receiver operating characteristics(AUROC)를 사용했다. IP 분류의 타당성을 검증하기 위해 TCGA cohort(n=110)에서 pathologist consensus ground-truth와의 상관관계를 Spearman correlation coefficient로 분석하였다. Clinical outcome 분석에서는 Kaplan-Meier method를 사용하여 PFS와 OS를 추정하고, Cox proportional hazards model을 통해 hazard ratios(HR)와 95% confidence intervals(CI)를 계산했다. 그룹 간 차이는 log-rank test로 평가했으며, categorical variables는 Fisher’s exact test, continuous variables는 Mann-Whitney U test로 비교했다. 다변량 분석은 multivariable logistic regression analysis를 통해 수행되었다.
강점
이 연구의 가장 큰 강점은 대규모 H&E-stained WSI 데이터를 기반으로 AI 모델을 개발하고, 이를 두 개의 독립적인 clinical cohort에서 검증했다는 점이다. 특히, TIL의 spatial distribution을 정량화하여 IP를 분류하는 방법은 기존 PD-L1 TPS만으로는 설명되지 않는 ICI 치료 반응을 더 잘 예측할 수 있음을 시사한다. 또한, genomic data와 multiplex IHC 분석을 통해 IP의 biological validity를 입증함으로써, AI 기반 biomarker의 신뢰성을 높였다. 마지막으로, 이 연구는 AI 기술이 routine pathology practice에서 어떻게 적용될 수 있는지를 구체적으로 제시하여, future clinical application에 대한 실용적인 통찰을 제공한다.
한계
첫째, 이 연구는 retrospective observational study로 설계되어 있어, causal relationship을 확립하기에는 한계가 있다. 둘째, TIL density threshold($> 106/\text{mm}^2$ for CE, $> 357/\text{mm}^2$ for CS)가 arbitrary하게 설정되었을 가능성이 있으며, 이는 다른 cohort나 institution에서 재현될 때 문제가 될 수 있다. 셋째, PD-L1 TPS ≥50%患者的 비율이 historic data보다 높게 보고되어, selection bias의 가능성을 배제하기 어렵다. 넷째, AI 모델의 performance는 internal 및 external validation에서 우수했지만, real-world setting에서의 robustness는 추가적인 prospective study를 통해 검증되어야 한다.
해석
이 연구 결과는 AI 기반 spatial analysis가 PD-L1 TPS를 보완하는 biomarker로서 ICI 치료 반응을 예측할 수 있음을 보여준다. 특히, inflamed IP와 높은 tumor response rate 및 prolonged survival 간의 연관성은, TIL의 spatial distribution이 immune microenvironment의 상태를 반영한다는 점을 강조한다. 이러한 발견은 future clinical trial에서 AI 기반 biomarker를 통합하여 patient stratification을 개선하는 데 기여할 수 있다. 또한, 이 연구는 oncology와 imaging 분야 간의 cross-disciplinary collaboration의 중요성을 부각시키며, LLM Wiki의 관련 문헌들과 함께 AI-driven precision medicine의 발전 방향을 제시한다.