Prediction-model development and temporal validation for six-month cancer-specific mortality in de novo metastatic gastroesophageal junction adenocarcinoma: a SEER-based study
Original Article

Prediction-model development and temporal validation for six-month cancer-specific mortality in de novo metastatic gastroesophageal junction adenocarcinoma: a SEER-based study

Lunjian Xia1# ORCID logo, Ziao Wang2#, Xiaoqian Guo1, Xueting Men1, Weikun Fan1, Qifeng Lu1

1Department of Gastroenterology, Fuyang People’s Hospital Affiliated to Bengbu Medical University, Fuyang, China; 2Department of Gastroenterology, First Hospital of Yangtze University, Jingzhou, China

Contributions: (I) Conception and design: L Xia, Z Wang, Q Lu; (II) Administrative support: Q Lu; (III) Provision of study materials or patients: L Xia, Z Wang; (IV) Collection and assembly of data: L Xia, Z Wang, X Guo, X Men; (V) Data analysis and interpretation: L Xia, Z Wang, X Guo, W Fan; (VI) Manuscript writing: All authors; (VII) Final approval of manuscript: All authors.

#These authors contributed equally to this work and share first authorship.

Correspondence to: Qifeng Lu, MD. Department of Gastroenterology, Fuyang People’s Hospital Affiliated to Bengbu Medical University, No. 501 Sanqing Road, Yingzhou District, Fuyang 236000, China. Email: luqifeng2010@yeah.net.

Background: De novo metastatic gastroesophageal junction adenocarcinoma (GEJA) is associated with heterogeneous short-term outcomes. A six-month cancer-specific mortality (CSM) horizon is clinically relevant because it reflects early disease lethality, limited physiological reserve, and the period in which first-line treatment tolerance and timely palliative discussions are often determined. Existing prediction tools have largely addressed long-term survival, postoperative disease, or metastatic gastric cancer as a broad entity. This study developed and temporally validated an interpretable prediction model for six-month CSM in de novo metastatic GEJA.

Methods: Patients diagnosed in 2010–2021 were identified from the Surveillance, Epidemiology, and End Results (SEER) database. Eligible cases had C16.0 cardia primary site, adenocarcinoma-related histology, malignant behavior, positive histology, single or first primary malignancy, de novo metastatic disease, and evaluable survival and cause-of-death information. Candidate predictors included demographic characteristics, histology, grade, tumor size, organ-specific metastases, diagnosis year, and first-course treatment-status fields. Patients diagnosed in 2010–2017 formed the development cohort; those diagnosed in 2018–2021 formed the temporal validation cohort. The endpoint was SEER-defined cancer-specific death within six months after diagnosis. Candidate models included logistic regression, elastic net, random forest, extreme gradient boosting (XGBoost), gradient boosting machine, and multilayer perceptron. Discrimination, calibration, Brier score, decision curves, threshold-specific sensitivity and specificity, and SHapley Additive exPlanations (SHAP) were evaluated.

Results: The analytic cohort included 6,564 patients, with 4,315 in development and 2,249 in temporal validation. Median age was 64 years, 5,189 patients (79.1%) were male, and liver metastasis was present in 3,249 patients (49.5%). Six-month CSM occurred in 1,941 patients (45.0%) in development and 902 patients (40.1%) in temporal validation. The selected XGBoost model achieved an area under the curve (AUC) of 0.808 [95% confidence interval (CI), 0.795–0.821] in development and 0.786 (95% CI, 0.767–0.806) in temporal validation. In temporal validation, the Brier score was 0.173 versus a null Brier score of 0.240, the calibration slope was 1.130, and the integrated calibration index was 0.026. At the high-risk threshold, sensitivity was 56.7%, specificity was 90.2%, positive predictive value was 79.5%, and negative predictive value was 75.7%. Observed six-month CSM was 20.2%, 32.1%, and 79.5% in low-, intermediate-, and high-risk groups, respectively.

Conclusions: A SEER-based prediction model was developed and temporally validated for six-month CSM in de novo metastatic GEJA. The model provided moderate discrimination, acceptable calibration, and clinically distinct population-level risk strata. Because validation was limited to a later SEER cohort, the model should be externally validated and locally recalibrated before clinical workflow use, and it must not be interpreted as evidence of treatment efficacy.

Keywords: Gastroesophageal junction adenocarcinoma (GEJA); Surveillance, Epidemiology, and End Results (SEER); cancer-specific mortality (CSM); machine learning; temporal validation


Submitted May 11, 2026. Accepted for publication Jul 07, 2026. Published online Jul 24, 2026.

doi: 10.21037/jgo-2026-0513


Highlight box

Key findings

• In 6,564 patients with de novo metastatic gastroesophageal junction adenocarcinoma (GEJA), six-month cancer-specific mortality (CSM) occurred in 45.0% of the development cohort and 40.1% of the temporal validation cohort.

• The selected extreme gradient boosting (XGBoost) model had a temporal-validation area under the curve (AUC) of 0.786, Brier score of 0.173, and integrated calibration index of 0.026.

• At the high-risk threshold, specificity and positive predictive value were 90.2% and 79.5%, respectively, while the low/intermediate threshold provided higher sensitivity (76.5%).

What is known and what is new?

• Previous registry-based tools mainly addressed broader metastatic gastric cancer, postoperative cohorts, or longer-term survival.

• This study focused on de novo metastatic C16.0 GEJA, retained organ-specific metastatic phenotypes, and evaluated discrimination, calibration, decision curves, threshold-level classification, and SHapley Additive exPlanations (SHAP)-based explanation.

What is the implication, and what should change now?

• The model may support population-level early risk stratification, prognostic communication, and trial enrichment after external validation; it should not be used to infer treatment effects or directly assign treatment.


Introduction

Gastroesophageal and gastric malignancies remain major contributors to the global cancer burden. GLOBOCAN 2022 estimated nearly 20 million new cancer cases and 9.7 million cancer deaths worldwide, and cancers of the stomach and esophagus remained among the leading causes of cancer death (1). The global burden of esophageal adenocarcinoma has continued to rise in several populations, with projections suggesting further increases by 2040 (2). Population-based analyses have also reported persistent mortality from gastroesophageal junction and esophageal adenocarcinoma despite changes in incidence and treatment patterns (3).

Gastroesophageal junction adenocarcinoma (GEJA) occupies a clinically difficult boundary between distal esophageal and proximal gastric cancer. Surveillance, Epidemiology, and End Results (SEER)-based studies have characterized its clinicopathological patterns and survival, and a competing-risk model has been developed for esophagogastric junction adenocarcinoma (4,5). These studies are useful for general survival estimation, but they do not directly address a narrower decision-facing question at first metastatic presentation: which patients are most likely to die from cancer within the next few months.

Early mortality in metastatic GEJA is clinically important because it may reflect a combination of metastatic burden, aggressive histology, limited physiological reserve, treatment ineligibility, rapid clinical decline, and competing supportive-care needs. A six-month horizon is not intended to replace long-term survival prediction. It captures an early interval during which first-line systemic therapy is commonly initiated or withheld, adverse reserve becomes clinically apparent, and early palliative-care discussions may be particularly relevant. This endpoint therefore has value for early prognostic communication, trial enrichment, and stratified observational research, provided it is not misread as a treatment-decision rule.

Several nomograms have been proposed for early death or survival in stage IV or metastatic gastric cancer (6-8). Other models have focused on multi-organ metastases or liver-metastatic gastric cancer (9,10). These studies support the feasibility of registry-based early-death prediction, but they usually pool gastric subsites or metastatic patterns. A model specific to de novo metastatic GEJA may be more anatomically coherent, because GEJA differs from distal gastric cancer in epidemiology, treatment pathways, and clinical framing at presentation.

Treatment for gastroesophageal and gastric adenocarcinoma has evolved substantially. Perioperative chemotherapy, neoadjuvant chemoradiotherapy, and adjuvant nivolumab have reshaped care for localized or resectable disease (11-15). For advanced gastric or GEJA, trastuzumab establishes human epidermal growth factor receptor 2 (HER2)-directed therapy (16), immune checkpoint inhibitor combinations improve first-line treatment in selected or broad populations (17-20), trastuzumab deruxtecan changes later-line HER2-positive disease (21), CLDN18.2-directed therapy extends biomarker-guided treatment (22,23), ramucirumab-based therapy remains an important later-line option (24,25), and recent regulatory updates further emphasize the rapidly changing treatment context (26). These advances reinforce the need for early risk assessment, but registry variables cannot measure performance status, biomarker status, treatment line, dose intensity, or response in sufficient detail to support causal treatment inference.

Machine-learning models have recently been applied to gastric cancer metastasis and survival prediction (27,28). For clinical prediction, however, model development should not stop at apparent discrimination. Transparent reporting, calibration assessment, threshold-level classification, decision curve analysis, and interpretable variable-attribution methods are needed before a model can be considered reproducible or clinically meaningful (29-32). Because the best functional form was uncertain, this study compared model families with different assumptions, including conventional logistic regression, penalized regression, tree-based bagging, gradient boosting, and neural-network modelling. The objective was to develop and temporally validate an interpretable SEER-based prediction model for six-month cancer-specific mortality (CSM) in de novo metastatic GEJA using routinely available registry variables and organ-specific metastatic phenotypes. We present this article in accordance with the TRIPOD reporting checklist (available at https://jgo.amegroups.com/article/view/10.21037/jgo-2026-0513/rc).


Methods

Study design and data source

This was a population-based, retrospective, multivariable prediction-model development and temporal validation study. Data were obtained from the SEER research database using SEER*Stat case-listing export. SEER is a population-based cancer registry resource with standardized tumor and survival fields, but it remains a registry rather than a clinical-trial or hospital electronic health record dataset. Its sampling frame supports population-level analyses, while its limitations include restricted variable granularity, possible coding transitions across calendar years, unmeasured confounding, incomplete treatment detail, and potential selection or reporting bias. This study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments.

Study population and disease definition

Patients diagnosed from 2010 to 2021 were eligible. The case definition required C16.0-cardia, not otherwise specified (NOS) primary site, malignant behavior, positive histology, known age, and adenocarcinoma-related histology. The positive-histology requirement provided histopathologic confirmation. Autopsy-only and death-certificate-only reporting sources were excluded. To reduce ambiguity from multiple primaries, the cohort was restricted to a single primary or first primary malignancy according to SEER sequence-number fields.

De novo metastatic disease was defined using the best available metastasis evidence across SEER staging eras, including organ-specific metastasis fields, derived M-stage, extent-of-disease metastasis recode, and collaborative stage metastasis variables. Liver, lung, bone, and brain metastases were retained as separate predictors rather than collapsed into a metastatic count, because different metastatic sites may carry different implications for early mortality. Treatment variables were extracted from first-course treatment fields and used only as prognostic adjustment variables, not as causal treatment exposures.

Outcome definition and follow-up handling

The primary endpoint was CSM within six months after diagnosis. Cancer-specific death was determined from the SEER cause-specific death classification and survival months. Patients who survived beyond six months, died of non-cancer causes within six months, or were censored before the six-month endpoint were classified as event-free for six-month CSM. Patients without evaluable survival time or cause-of-death classification were excluded from the final analytic cohort. The endpoint should be interpreted with the usual caution for registry cause-of-death coding, because misclassification is possible in patients with comorbidities or competing terminal events.

Candidate predictors and preprocessing

Candidate predictors were selected a priori according to clinical relevance, SEER availability, cross-year definitional stability, and plausible availability at or near initial metastatic diagnosis. They included age at diagnosis, sex, race and ethnicity, marital status, year of diagnosis, histologic subtype, tumor grade, tumor size, liver metastasis, lung metastasis, bone metastasis, brain metastasis, primary-site surgery, radiotherapy, and chemotherapy. Variables with non-harmonizable definitions or insufficient usable information across the study period were not retained.

Tumor size was harmonized from collaborative stage tumor size for earlier diagnosis years and tumor size summary for later diagnosis years. Tumor grade was harmonized when available, while unknown grade was retained as a separate category. Missing proportions for candidate predictors are reported in Table 1. For model fitting, missing tumor size values were replaced with the development-cohort median and accompanied by an explicit tumor-size-missingness indicator. Unknown categories were retained when they reflected registry coding or clinically meaningful absence of information. Structural missingness in tumor grade, which was unavailable in the 2018-2021 temporal validation cohort because of SEER coding changes, was not imputed.

Table 1

Baseline characteristics of patients in the development and temporal validation cohorts

Characteristic Overall (n=6,564) Development (n=4,315) Temporal validation (n=2,249) P value
Age, years 64.0 (56.0–72.0) 64.0 (56.0–72.0) 65.0 (57.0–73.0) 0.009
Tumor size, mm 47.0 (30.0–61.0) 46.0 (30.0–60.0) 49.5 (30.0–63.0) 0.19
Sex 0.28
   Male 5,189 (79.1) 3,428 (79.4) 1,761 (78.3)
   Female 1,375 (20.9) 887 (20.6) 488 (21.7)
Race/ethnicity <0.001
   Non-Hispanic White 4,762 (72.5) 3,206 (74.3) 1,556 (69.2)
   Hispanic 872 (13.3) 535 (12.4) 337 (15.0)
   Non-Hispanic Asian or Pacific Islander 464 (7.1) 284 (6.6) 180 (8.0)
   Non-Hispanic Black 407 (6.2) 253 (5.9) 154 (6.8)
   Other/unknown 59 (0.9) 37 (0.9) 22 (1.0)
Marital status 0.001
   Married 3,866 (58.9) 2,571 (59.6) 1,295 (57.6)
   Unmarried 2,427 (37.0) 1,544 (35.8) 883 (39.3)
   Unknown 271 (4.1) 200 (4.6) 71 (3.2)
Histologic subtype <0.001
   Adenocarcinoma, NOS 5,276 (80.4) 3,395 (78.7) 1,881 (83.6)
   Signet ring cell carcinoma 695 (10.6) 516 (12.0) 179 (8.0)
   Intestinal-type adenocarcinoma 292 (4.4) 191 (4.4) 101 (4.5)
   Mucinous adenocarcinoma 102 (1.6) 79 (1.8) 23 (1.0)
   Diffuse-type adenocarcinoma 97 (1.5) 55 (1.3) 42 (1.9)
   Other adenocarcinoma 79 (1.2) 59 (1.4) 20 (0.9)
   Tubular/papillary adenocarcinoma 23 (0.4) 20 (0.5) 3 (0.1)
Tumor grade <0.001
   Unknown 2,976 (45.3) 727 (16.8) 2,249 (100.0)
   Poorly differentiated 2,337 (35.6) 2,337 (54.2) 0 (0.0)
   Moderately differentiated 1,149 (17.5) 1,149 (26.6) 0 (0.0)
   Well differentiated 102 (1.6) 102 (2.4) 0 (0.0)
Tumor size missingness 0.95
   Missing 3,420 (52.1) 2,247 (52.1) 1,173 (52.2)
   Available 3,144 (47.9) 2,068 (47.9) 1,076 (47.8)
Bone metastasis <0.001
   No 5,429 (82.7) 3,551 (82.3) 1,878 (83.5)
   Yes 947 (14.4) 604 (14.0) 343 (15.3)
   Unknown 188 (2.9) 160 (3.7) 28 (1.2)
Brain metastasis <0.001
   No 6,172 (94.0) 4,011 (93.0) 2,161 (96.1)
   Yes 177 (2.7) 126 (2.9) 51 (2.3)
   Unknown 215 (3.3) 178 (4.1) 37 (1.6)
Liver metastasis <0.001
   No 3,177 (48.4) 2,113 (49.0) 1,064 (47.3)
   Yes 3,249 (49.5) 2,083 (48.3) 1,166 (51.8)
   Unknown 138 (2.1) 119 (2.8) 19 (0.8)
Lung metastasis <0.001
   No 5,192 (79.1) 3,351 (77.7) 1,841 (81.9)
   Yes 1,136 (17.3) 767 (17.8) 369 (16.4)
   Unknown 236 (3.6) 197 (4.6) 39 (1.7)
Primary-site surgery 0.01
   No surgery 5,978 (91.1) 3,900 (90.4) 2,078 (92.4)
   Primary-site surgery 572 (8.7) 403 (9.3) 169 (7.5)
   Unknown 14 (0.2) 12 (0.3) 2 (0.1)
Radiotherapy <0.001
   No/unknown radiation 4,574 (69.7) 2,941 (68.2) 1,633 (72.6)
   Radiation 1,986 (30.3) 1,371 (31.8) 615 (27.3)
   Unknown 4 (0.1) 3 (0.1) 1 (0.0)
Chemotherapy 0.09
   Chemotherapy 4,676 (71.2) 3,044 (70.5) 1,632 (72.6)
   No/unknown chemotherapy 1,888 (28.8) 1,271 (29.5) 617 (27.4)
6-month cancer-specific death <0.001
   No 3,721 (56.7) 2,374 (55.0) 1,347 (59.9)
   Yes 2,843 (43.3) 1,941 (45.0) 902 (40.1)
6-month other-cause death 0.35
   No 6,380 (97.2) 4,200 (97.3) 2,180 (96.9)
   Yes 184 (2.8) 115 (2.7) 69 (3.1)

Continuous variables are reported as median (interquartile range), and categorical variables are reported as n (%). P values compare development and temporal validation cohorts. Tumor grade was missing for all patients in the temporal validation cohort [2018–2021] due to changes in SEER data collection and cross-year coding availability. This limitation is addressed in the Discussion section. NOS, not otherwise specified; SEER, Surveillance, Epidemiology, and End Results.

Age and tumor size were represented flexibly using spline terms where appropriate. Categorical variables were encoded as indicator variables for model fitting. SEER chemotherapy is recorded as yes versus no/unknown; therefore, the variable was retained as its original registry treatment-status field and was not separated into true untreated and unknown categories.

Development, temporal validation, and model evaluation

Patients diagnosed in 2010–2017 formed the development cohort, and those diagnosed in 2018–2021 formed the temporal validation cohort. Within the development cohort, model fitting and Platt calibration were performed using an internal training-calibration split. The temporal validation cohort was held out from model fitting, calibration, and tuning, and was used only for final evaluation.

Representative model families were chosen to reflect different modelling assumptions: logistic regression as a conventional linear baseline, elastic net as a penalized regression method, random forest as a bagging-based tree ensemble, extreme gradient boosting (XGBoost) and gradient boosting machine as boosting-based tree ensembles, and multilayer perceptron as a neural-network classifier. Model selection prioritized the full validation profile, including temporal-validation discrimination, calibration, Brier score, decision curve analysis, threshold-level classification, and interpretability, rather than development-set area under the curve (AUC) alone.

AUC values with 95% confidence intervals (CIs) were estimated using DeLong’s method. Calibration was assessed using calibration intercept, calibration slope, Brier score, null Brier score, integrated calibration index, and calibration plots. Decision curve analysis was used to evaluate net benefit across clinically plausible threshold probabilities (31). SHapley Additive exPlanations (SHAP) analysis was used to explain the final tree-based model at the feature level (32).

No single AUC value was treated as sufficient to justify clinical use. In this oncology setting, temporal-validation AUC around 0.80 was considered compatible with population-level risk stratification only if calibration, threshold-level classification, and decision curves were acceptable. It was not considered sufficient for direct treatment assignment. Two prespecified risk cutoffs derived from tertiles of development-cohort predictions were applied unchanged to the temporal validation cohort: 26.5% for the low/intermediate boundary and 40.5% for the intermediate/high boundary. Sensitivity, specificity, positive predictive value, negative predictive value, and accuracy were calculated for these thresholds in temporal validation.

Sensitivity analyses

Prespecified sensitivity analyses evaluated a no-treatment model, the locked six-month risk score applied to 12-month CSM, candidate-model ranking, risk-group stability, and feature-contribution summaries. The no-treatment model removed primary-site surgery, radiotherapy, and chemotherapy to assess the model’s dependence on first-course treatment-status fields. Because SEER chemotherapy combines no and unknown chemotherapy, and because grade availability changed after 2018, both issues were addressed explicitly in interpretation rather than treated as causal findings.

Statistical analysis

Continuous variables are summarized as median (interquartile range) and were compared between the development and temporal validation cohorts using the Wilcoxon rank-sum test. Categorical variables are summarized as number (percentage) and were compared using Pearson's chi-square test; when expected cell counts were <5, Monte Carlo simulated chi-square P values based on 2,000 replicates were used. AUCs and 95% confidence intervals were estimated using DeLong’s method. All statistical tests were two-sided, and P<0.05 was considered statistically significant. Analyses were performed in R.


Results

Cohort derivation and baseline characteristics

The SEER*Stat export identified 22,772 patients with C16.0 cardia primary site, positive histology, malignant behavior, and diagnosis from 2010 to 2021. After restriction to adenocarcinoma-related histology, single or first primary malignancy, de novo metastatic disease, and evaluable six-month CSM status, 6,564 patients remained in the analytic cohort. The development cohort included 4,315 patients diagnosed in 2010–2017, and the temporal validation cohort included 2,249 patients diagnosed in 2018–2021 (Figure 1).

Figure 1 Cohort selection and model development workflow. (A) Shows identification of patients with de novo metastatic gastroesophageal junction adenocarcinoma from the SEER database. The final analytic cohort included 6,564 patients and was divided into a development cohort diagnosed in 2010–2017 (n=4,315) and a temporal validation cohort diagnosed in 2018–2021 (n=2,249). (B) Shows the locked modelling workflow. Candidate models were trained and calibrated within the development cohort only; the temporal validation cohort was used only for final model evaluation. AUC, area under the curve; CS, collaborative staging; CSM, cancer-specific mortality; DCA, decision curve analysis; EOD, extent of disease; GBM, gradient boosting machine; ICI, integrated calibration index; MLP, multilayer perceptron; NOS, not otherwise specified; SEER, Surveillance, Epidemiology, and End Results; SHAP, SHapley Additive exPlanations; XGBoost, extreme gradient boosting.

The median age was 64 years in the full cohort, and 5,189 patients (79.1%) were male. Non-Hispanic White patients accounted for 4,762 cases (72.5%). Adenocarcinoma, NOS was the dominant histologic subtype (5,276 patients, 80.4%), followed by signet ring cell carcinoma (695 patients, 10.6%). Liver metastasis was the most common organ-specific metastatic phenotype (3,249 patients, 49.5%), followed by lung metastasis (1,136 patients, 17.3%) and bone metastasis (947 patients, 14.4%). Chemotherapy was recorded in 4,676 patients (71.2%). Tumor grade was unavailable for all patients in the temporal validation cohort because of SEER coding availability after 2018; this was retained as a database limitation rather than imputed (Table 1).

Six-month CSM

Six-month cancer-specific death occurred in 1,941 of 4,315 patients (45.0%) in the development cohort and 902 of 2,249 patients (40.1%) in the temporal validation cohort. Six-month other-cause death occurred in 115 patients (2.7%) in development and 69 patients (3.1%) in temporal validation.

Candidate model comparison and final model selection

Model performance is summarized in Table 2. Logistic regression, elastic net, GBM, and XGBoost produced closely similar temporal-validation AUCs. Random forest showed apparent overfitting, with a development AUC of 0.924 but a temporal-validation AUC of 0.767 and a development calibration slope of 1.835, and was therefore not selected despite high apparent discrimination. MLP also showed a larger decline from development to temporal validation. The selected XGBoost model was retained because it combined competitive temporal-validation discrimination, acceptable calibration, decision-curve usefulness, threshold-level separation, and direct SHAP-based interpretability (Figures 2,3).

Table 2

Performance of representative candidate models

Model family Model Dataset N Events Event rate, % AUC (95% CI) Brier score Null Brier Calibration intercept Calibration slope ICI Log loss
Logistic regression Logistic regression Development 4,315 1,941 45 0.797 (0.784–0.810) 0.179 0.247 0.177 1.03 0.03 0.534
Logistic regression Logistic regression Temporal validation 2,249 902 40.1 0.784 (0.764–0.803) 0.174 0.24 0.21 1.036 0.039 0.528
Elastic net ElasticNet_alpha_0.75 Development 4,315 1,941 45 0.796 (0.783–0.809) 0.18 0.247 0.168 1.011 0.032 0.536
Elastic net ElasticNet_alpha_0.75 Temporal validation 2,249 902 40.1 0.786 (0.767–0.806) 0.173 0.24 0.201 1.026 0.031 0.524
Random forest RandomForest_mtry_6 Development 4,315 1,941 45 0.924 (0.916–0.932) 0.126 0.247 0.563 1.835 0.119 0.405
Random forest RandomForest_mtry_6 Temporal validation 2,249 902 40.1 0.767 (0.746–0.787) 0.179 0.24 0.085 1.061 0.047 0.541
XGBoost XGBoost_6 Development 4,315 1,941 45 0.808 (0.795–0.821) 0.177 0.247 0.192 1.061 0.044 0.529
XGBoost XGBoost_6 Temporal validation 2,249 902 40.1 0.786 (0.767–0.806) 0.173 0.24 0.18 1.13 0.026 0.524
GBM GBM_2 Development 4,315 1,941 45 0.813 (0.801–0.826) 0.174 0.247 0.194 1.057 0.041 0.523
GBM GBM_2 Temporal validation 2,249 902 40.1 0.785 (0.765–0.804) 0.173 0.24 0.112 1.019 0.026 0.524
MLP MLP_size5_decay0.1 Development 4,315 1,941 45 0.831 (0.819–0.843) 0.168 0.247 0.297 1.366 0.057 0.51
MLP MLP_size5_decay0.1 Temporal validation 2,249 902 40.1 0.729 (0.708–0.751) 0.203 0.24 0.089 0.827 0.067 0.6

AUC values with 95% CIs were estimated using DeLong’s method. Calibration metrics were calculated after Platt calibration where applicable. Random forest displayed strong optimism in the development cohort and was therefore not considered for final deployment despite its high apparent development AUC. AUC, area under the receiver operating characteristic curve; CI, confidence interval; GBM, gradient boosting machine; ICI, integrated calibration index; MLP, multilayer perceptron; XGBoost, extreme gradient boosting.

Figure 2 Receiver operating characteristic curves of the final locked XGBoost model in the development cohort and temporal validation cohort. The diagonal dashed line indicates the reference line corresponding to an AUC of 0.5. AUC values with 95% confidence intervals were estimated using DeLong’s method. AUC, area under the receiver operating characteristic curve; CI, confidence interval; XGBoost, extreme gradient boosting.
Figure 3 Receiver operating characteristic curves of representative candidate models in the temporal validation cohort (n=2,249). One representative model from each model family is shown. AUC values and 95% confidence intervals were estimated using DeLong’s method; point estimates are shown in the legend. AUC, area under the receiver operating characteristic curve; GBM, gradient boosting machine; MLP, multilayer perceptron; XGBoost, extreme gradient boosting.

Final model performance, threshold behavior, and risk stratification

The selected XGBoost model achieved an AUC of 0.808 (95% CI, 0.795–0.821) in the development cohort and 0.786 (95% CI, 0.767–0.806) in the temporal validation cohort. In temporal validation, the Brier score was lower than the null Brier score (0.173 vs. 0.240), the calibration intercept was 0.180, the calibration slope was 1.130, and the integrated calibration index was 0.026 (Table 3).

Table 3

Final selected model and no-treatment sensitivity analysis

Analysis Dataset Model N Events Event rate, % AUC (95% CI) Brier score Null Brier Calibration intercept Calibration slope ICI
Final selected model Development XGBoost_6 4,315 1,941 45 0.808 (0.795–0.821) 0.177 0.247 0.192 1.061 0.044
Final selected model Temporal validation XGBoost_6 2,249 902 40.1 0.786 (0.767–0.806) 0.173 0.24 0.18 1.13 0.026
No-treatment sensitivity model Development No-treatment sensitivity model 4,315 1,941 45 0.683 (0.667–0.699) 0.224 0.247 0.159 1.073 0.033
No-treatment sensitivity model Temporal validation No-treatment sensitivity model 2,249 902 40.1 0.650 (0.628–0.673) 0.227 0.24 0.101 0.859 0.051

Treatment variables were used for prognostic adjustment, not causal inference. The no-treatment model intentionally removed primary-site surgery, radiotherapy, and chemotherapy variables to evaluate dependence on first-course treatment fields. AUC, area under the receiver operating characteristic curve; CI, confidence interval; ICI, integrated calibration index; XGBoost, extreme gradient boosting.

The threshold-specific classification profile is shown in Table 4. Using the 26.5% low/intermediate cutoff, classification of intermediate- or high-risk patients as test-positive yielded sensitivity of 76.5%, specificity of 62.1%, positive predictive value of 57.5%, and negative predictive value of 79.8%. Using the 40.5% intermediate/high cutoff, classification of high-risk patients as test-positive yielded lower sensitivity (56.7%) but higher specificity (90.2%) and positive predictive value (79.5%). This pattern supports use of the model for risk stratification rather than binary treatment assignment.

Table 4

Threshold-specific classification performance of the final model in the temporal validation cohort

Threshold Predicted risk cutoff Positive classification True positives False positives Sensitivity, % Specificity, % PPV, % NPV, % Accuracy, %
Low/intermediate threshold ≥26.5% Intermediate or high risk 690 510 76.5 62.1 57.5 79.8 67.9
Intermediate/high threshold ≥40.5% High risk 511 132 56.7 90.2 79.5 75.7 76.7

Risk thresholds were derived from tertiles of final-model predictions in the development cohort and applied unchanged to temporal validation. NPV, negative predictive value; PPV, positive predictive value.

The locked risk groups separated observed six-month CSM in temporal validation. Observed event rates were 20.2% in the low-risk group, 32.1% in the intermediate-risk group, and 79.5% in the high-risk group. The mean predicted risks in these groups were 0.2007, 0.3101, and 0.7371, respectively. Calibration, clinical impact, and decision curve analyses are shown in Figure 4.

Figure 4 Calibration, clinical impact, and decision curve analyses for the final locked XGBoost model in the temporal validation cohort (n=2,249). (A) Shows decile-based calibration with 95% confidence intervals and the 45-degree ideal calibration line. (B) Shows the number of patients classified as high risk and the number of high-risk patients with observed six-month cancer-specific death per 1,000 patients across threshold probabilities. (C) Shows decision curve analysis comparing the model with treat-all and treat-none strategies. XGBoost, extreme gradient boosting.

Sensitivity analyses and model explanation

Removing first-course treatment variables reduced temporal-validation discrimination from 0.786 to 0.650, indicating that treatment-status fields carried substantial prognostic information but also confirming that the final model relied partly on registry treatment variables (Table 3). The 12-month sensitivity analysis showed that the locked six-month risk score remained associated with longer-horizon CSM but performed less strongly for 12-month prediction, with a temporal-validation AUC of 0.732. This supported the intended use of the model as an early six-month risk tool rather than a general long-term survival model.

SHAP analysis identified chemotherapy status, primary-site surgery status, liver metastasis, tumor grade, age, bone metastasis, diagnosis year, lung metastasis, signet ring cell histology, marital status, radiotherapy, tumor size, and brain metastasis as influential predictors (Figure 5). The large contributions of chemotherapy and surgery fields should be interpreted as prognostic information captured in registry treatment-status variables, not as evidence of causal treatment benefit or harm.

Figure 5 SHAP summary plot of the final XGBoost model. Each point represents one patient-level contribution for a given feature. The horizontal axis shows the SHAP value; positive SHAP values increase the predicted probability of six-month cancer-specific death and negative values decrease it. The color scale represents relative feature value, with red indicating higher values and blue indicating lower values. Treatment variables were used for prognostic adjustment and should not be interpreted causally. SHAP, SHapley Additive exPlanations; XGBoost, extreme gradient boosting.

Discussion

Principal findings

In this population-based study, an interpretable machine-learning model was developed and temporally validated to estimate six-month CSM in patients with de novo metastatic GEJA. The model achieved moderate temporal-validation discrimination, acceptable calibration, a Brier score lower than the null Brier score, and clearly separated risk strata. The high-risk group had an observed six-month CSM rate of nearly 80%, whereas the low-risk group had an observed rate of approximately 20%. These results suggest that registry variables can stratify early CSM risk at the population level when model development is accompanied by chronological validation, calibration assessment, threshold-level evaluation, and cautious interpretation.

Comparison with previous models

The study addresses a distinct clinical gap. Prior SEER-based work in esophagogastric junction adenocarcinoma has focused largely on overall survival, competing-risk survival, or broad clinicopathologic description (4,5). Early-death models in metastatic gastric cancer have shown that registry variables can identify high-risk patients, but those studies generally pooled gastric cancer subsites or focused on specific metastatic patterns such as liver metastasis or multi-organ spread (6-10). By restricting the cohort to C16.0 GEJA, requiring de novo metastatic disease, and modelling liver, lung, bone, and brain metastases separately, the present study avoids a generic metastatic-count approach and remains closer to the clinical anatomy of GEJA.

The model-selection results also show why algorithm ranking by apparent AUC alone is inadequate. Random forest had a high development AUC but weaker temporal validation and marked calibration optimism, a pattern compatible with overfitting. XGBoost was selected because it balanced discrimination, calibration, Brier score, decision-curve performance, threshold behavior, and interpretability. The similar temporal-validation AUCs of elastic net, GBM, and XGBoost indicate that the endpoint signal was not confined to one algorithm; the final selection was based on the overall validation profile rather than a marginal AUC difference.

Clinical interpretation

The six-month horizon should be interpreted as an early prognosis window, not as a surrogate for treatment effect. In metastatic GEJA, early death may arise from extensive metastatic burden, aggressive histology, comorbidity, frailty, rapidly progressive symptoms, limited access to systemic therapy, or inability to tolerate treatment. The model can identify groups with materially different early mortality risk, which may be useful for risk communication, observational stratification, and trial enrichment. However, it should not be used to deny or assign systemic therapy, surgery, radiotherapy, or palliative-care interventions.

The threshold analysis provides a more clinically interpretable view than AUC alone. The low/intermediate cutoff favored sensitivity and may be more useful for identifying a broad group requiring closer early assessment. The intermediate/high cutoff favored specificity and positive predictive value, identifying a smaller subgroup with very high observed early CSM. These trade-offs reinforce that model use depends on the clinical purpose. A threshold that is useful for screening or research enrichment is not necessarily appropriate for high-stakes treatment decisions.

The SHAP analysis supported several clinically plausible patterns, including contributions from organ-specific metastases, age, histology, grade, and registry treatment-status fields. The treatment-related findings require particular restraint. SEER chemotherapy is coded as yes versus no/unknown. The no/unknown group may include patients who were too ill for chemotherapy, declined treatment, received treatment outside reporting pathways, died rapidly before treatment, or had incomplete documentation. Similarly, primary-site surgery status may reflect selection and documentation processes. These variables were retained for prognostic adjustment but cannot be interpreted causally.

Limitations

The most important limitation is the absence of validation outside SEER. Temporal validation is stricter than a random split because it tests performance across calendar time and coding transitions, but it remains internal to the same registry system, with the same broad data architecture, variable definitions, and reporting practices. It does not establish transportability to other registries, countries, hospital networks, or treatment eras. Independent validation in non-SEER datasets and local recalibration would be required before clinical workflow use.

Several registry limitations also affect interpretation. SEER lacks Eastern Cooperative Oncology Group performance status, laboratory values, nutritional status, liver and renal function, inflammatory markers, detailed tumor burden, HER2 status, programmed death-ligand 1 expression, microsatellite instability, CLDN18.2 expression, treatment line, chemotherapy regimen, immunotherapy exposure, dose intensity, response, and supportive-care details. Cause-specific death classification may be imperfect, especially in patients with comorbidities or competing terminal events. Tumor grade was unavailable for all temporal-validation patients because of SEER coding changes after 2018. This missingness was treated as structural database missingness rather than imputed.

The study should therefore be understood as a population-based early-risk stratification analysis, not as a bedside treatment-assignment instrument. A future model that combines SEER-like population coverage with prospective clinical variables, biomarker data, treatment sequence, and external validation could improve calibration, clinical interpretability, and transportability.


Conclusions

A SEER-based XGBoost model was developed and temporally validated to estimate six-month CSM in de novo metastatic GEJA. The model provided moderate discrimination, acceptable calibration, clinically distinct risk strata, and interpretable feature attribution. Its appropriate role is early prognostic risk stratification and research support after external validation and local recalibration, not causal treatment inference or direct treatment assignment.


Acknowledgments

The authors thank the SEER Program registries for providing access to population-based cancer data.


Footnote

Reporting Checklist: The authors have completed the TRIPOD reporting checklist. Available at https://jgo.amegroups.com/article/view/10.21037/jgo-2026-0513/rc

Peer Review File: Available at https://jgo.amegroups.com/article/view/10.21037/jgo-2026-0513/prf

Funding: This study was supported by the 12th Batch of Municipal-Level Industrial Innovation Team Projects in Fuyang City and the 2025 Research Project for the Joint Development Programme for Physicians Specialising in Early-Stage Gastrointestinal Cancer (No. GTCZ-2025-23).

Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://jgo.amegroups.com/article/view/10.21037/jgo-2026-0513/coif). The authors have no conflicts of interest to declare.

Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. This study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments.

Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.


References

  1. Bray F, Laversanne M, Sung H, et al. Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA Cancer J Clin 2024;74:229-63. [Crossref] [PubMed]
  2. Morgan E, Soerjomataram I, Rumgay H, et al. The Global Landscape of Esophageal Squamous Cell Carcinoma and Esophageal Adenocarcinoma Incidence and Mortality in 2020 and Projections to 2040: New Estimates From GLOBOCAN 2020. Gastroenterology 2022;163:649-658.e2. [Crossref] [PubMed]
  3. Agarwal S, Bell MG, Dhaliwal L, et al. Population Based Time Trends in the Epidemiology and Mortality of Gastroesophageal Junction and Esophageal Adenocarcinoma. Dig Dis Sci 2024;69:246-53. [Crossref] [PubMed]
  4. Liu X, Jiang Q, Yue C, et al. Clinicopathological Characteristics and Survival Predictions for Adenocarcinoma of the Esophagogastric Junction: A SEER Population-Based Retrospective Study. Int J Gen Med 2021;14:10303-14. [Crossref] [PubMed]
  5. Wang T, Wu Y, Zhou H, et al. Development and validation of a novel competing risk model for predicting survival of esophagogastric junction adenocarcinoma: a SEER population-based study and external validation. BMC Gastroenterol 2021;21:38. [Crossref] [PubMed]
  6. Zhu Y, Fang X, Wang L, et al. A Predictive Nomogram for Early Death of Metastatic Gastric Cancer: A Retrospective Study in the SEER Database and China. J Cancer 2020;11:5527-35. [Crossref] [PubMed]
  7. Feng Y, Guo K, Jin H, et al. A Predictive Nomogram for Early Mortality in Stage IV Gastric Cancer. Med Sci Monit 2020;26:e923931. [Crossref] [PubMed]
  8. Yang Y, Chen ZJ, Yan S. The incidence, risk factors and predictive nomograms for early death among patients with stage IV gastric cancer: a population-based study. J Gastrointest Oncol 2020;11:964-82. [Crossref] [PubMed]
  9. Wang T, Liu C, Wang W, et al. Development and validation of a Surveillance, Epidemiology, and End Results (SEER)-based prognostic nomogram for predicting survival in gastric cancer with multi-organ metastases. Transl Cancer Res 2022;11:1534-51. [Crossref] [PubMed]
  10. Meng N, Niu X, Wu J, et al. Development and validation of nomogram models for predicting overall survival and cancer-specific survival in gastric cancer patients with liver metastases: a cohort study based on the SEER database. Am J Cancer Res 2024;14:2272-86. [Crossref] [PubMed]
  11. Al-Batran SE, Homann N, Pauligk C, et al. Perioperative chemotherapy with fluorouracil plus leucovorin, oxaliplatin, and docetaxel versus fluorouracil or capecitabine plus cisplatin and epirubicin for locally advanced, resectable gastric or gastro-oesophageal junction adenocarcinoma (FLOT4): a randomised, phase 2/3 trial. Lancet 2019;393:1948-57. [Crossref] [PubMed]
  12. Cunningham D, Allum WH, Stenning SP, et al. Perioperative chemotherapy versus surgery alone for resectable gastroesophageal cancer. N Engl J Med 2006;355:11-20. [Crossref] [PubMed]
  13. van Hagen P, Hulshof MC, van Lanschot JJ, et al. Preoperative chemoradiotherapy for esophageal or junctional cancer. N Engl J Med 2012;366:2074-84. [Crossref] [PubMed]
  14. Shapiro J, van Lanschot JJB, Hulshof MCCM, et al. Neoadjuvant chemoradiotherapy plus surgery versus surgery alone for oesophageal or junctional cancer (CROSS): long-term results of a randomised controlled trial. Lancet Oncol 2015;16:1090-8. [Crossref] [PubMed]
  15. Kelly RJ, Ajani JA, Kuzdzal J, et al. Adjuvant Nivolumab in Resected Esophageal or Gastroesophageal Junction Cancer. N Engl J Med 2021;384:1191-203. [Crossref] [PubMed]
  16. Bang YJ, Van Cutsem E, Feyereislova A, et al. Trastuzumab in combination with chemotherapy versus chemotherapy alone for treatment of HER2-positive advanced gastric or gastro-oesophageal junction cancer (ToGA): a phase 3, open-label, randomised controlled trial. Lancet 2010;376:687-97. [Crossref] [PubMed]
  17. Janjigian YY, Shitara K, Moehler M, et al. First-line nivolumab plus chemotherapy versus chemotherapy alone for advanced gastric, gastro-oesophageal junction, and oesophageal adenocarcinoma (CheckMate 649): a randomised, open-label, phase 3 trial. Lancet 2021;398:27-40. [Crossref] [PubMed]
  18. Kang YK, Chen LT, Ryu MH, et al. Nivolumab plus chemotherapy versus placebo plus chemotherapy in patients with HER2-negative, untreated, unresectable advanced or recurrent gastric or gastro-oesophageal junction cancer (ATTRACTION-4): a randomised, multicentre, double-blind, placebo-controlled, phase 3 trial. Lancet Oncol 2022;23:234-47. [Crossref] [PubMed]
  19. Rha SY, Oh DY, Yañez P, et al. Pembrolizumab plus chemotherapy versus placebo plus chemotherapy for HER2-negative advanced gastric cancer (KEYNOTE-859): a multicentre, randomised, double-blind, phase 3 trial. Lancet Oncol 2023;24:1181-95. [Crossref] [PubMed]
  20. Janjigian YY, Kawazoe A, Bai Y, et al. Pembrolizumab plus trastuzumab and chemotherapy for HER2-positive gastric or gastro-oesophageal junction adenocarcinoma: interim analyses from the phase 3 KEYNOTE-811 randomised placebo-controlled trial. Lancet 2023;402:2197-208. [Crossref] [PubMed]
  21. Shitara K, Bang YJ, Iwasa S, et al. Trastuzumab Deruxtecan in Previously Treated HER2-Positive Gastric Cancer. N Engl J Med 2020;382:2419-30. [Crossref] [PubMed]
  22. Shitara K, Lordick F, Bang YJ, et al. Zolbetuximab plus mFOLFOX6 in patients with CLDN18.2-positive, HER2-negative, untreated, locally advanced unresectable or metastatic gastric or gastro-oesophageal junction adenocarcinoma (SPOTLIGHT): a multicentre, randomised, double-blind, phase 3 trial. Lancet 2023;401:1655-68. [Crossref] [PubMed]
  23. Shah MA, Shitara K, Ajani JA, et al. Zolbetuximab plus CAPOX in CLDN18.2-positive gastric or gastroesophageal junction adenocarcinoma: the randomized, phase 3 GLOW trial. Nat Med 2023;29:2133-41. [Crossref] [PubMed]
  24. Fuchs CS, Tomasek J, Yong CJ, et al. Ramucirumab monotherapy for previously treated advanced gastric or gastro-oesophageal junction adenocarcinoma (REGARD): an international, randomised, multicentre, placebo-controlled, phase 3 trial. Lancet 2014;383:31-9. [Crossref] [PubMed]
  25. Wilke H, Muro K, Van Cutsem E, et al. Ramucirumab plus paclitaxel versus placebo plus paclitaxel in patients with previously treated advanced gastric or gastro-oesophageal junction adenocarcinoma (RAINBOW): a double-blind, randomised phase 3 trial. Lancet Oncol 2014;15:1224-35. [Crossref] [PubMed]
  26. U.S. Food and Drug Administration. FDA approves pembrolizumab for HER2 positive gastric or gastroesophageal junction adenocarcinoma expressing PD-L1 (CPS >=1). Published March 19, 2025. Accessed June 30, 2026. Available online: https://www.fda.gov/drugs/resources-information-approved-drugs/fda-approves-pembrolizumab-her2-positive-gastric-or-gastroesophageal-junction-adenocarcinoma
  27. Qin X, Qiu B, Ge L, et al. Applying machine learning techniques to predict the risk of distant metastasis from gastric cancer: a real world retrospective study. Front Oncol 2024;14:1455914. [Crossref] [PubMed]
  28. Zhang C, Zhang Y, Yang YH, et al. Machine learning models for predicting one-year survival in patients with metastatic gastric cancer who experienced upfront radical gastrectomy. Front Mol Biosci 2022;9:937242. [Crossref] [PubMed]
  29. Collins GS, Reitsma JB, Altman DG, et al. Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (TRIPOD): the TRIPOD Statement. BMC Med 2015;13:1. [Crossref] [PubMed]
  30. Collins GS, Moons KGM, Dhiman P, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ 2024;385:e078378. [Crossref] [PubMed]
  31. Vickers AJ, Elkin EB. Decision curve analysis: a novel method for evaluating prediction models. Med Decis Making 2006;26:565-74. [Crossref] [PubMed]
  32. Lundberg SM, Erion G, Chen H, et al. From Local Explanations to Global Understanding with Explainable AI for Trees. Nat Mach Intell 2020;2:56-67. [Crossref] [PubMed]
Cite this article as: Xia L, Wang Z, Guo X, Men X, Fan W, Lu Q. Prediction-model development and temporal validation for six-month cancer-specific mortality in de novo metastatic gastroesophageal junction adenocarcinoma: a SEER-based study. J Gastrointest Oncol 2026;17(4):209. doi: 10.21037/jgo-2026-0513

Download Citation