Prediction-model development and temporal validation for six-month cancer-specific mortality in de novo metastatic gastroesophageal junction adenocarcinoma: a SEER-based study
Highlight box
Key findings
• In 6,564 patients with de novo metastatic gastroesophageal junction adenocarcinoma (GEJA), six-month cancer-specific mortality (CSM) occurred in 45.0% of the development cohort and 40.1% of the temporal validation cohort.
• The selected extreme gradient boosting (XGBoost) model had a temporal-validation area under the curve (AUC) of 0.786, Brier score of 0.173, and integrated calibration index of 0.026.
• At the high-risk threshold, specificity and positive predictive value were 90.2% and 79.5%, respectively, while the low/intermediate threshold provided higher sensitivity (76.5%).
What is known and what is new?
• Previous registry-based tools mainly addressed broader metastatic gastric cancer, postoperative cohorts, or longer-term survival.
• This study focused on de novo metastatic C16.0 GEJA, retained organ-specific metastatic phenotypes, and evaluated discrimination, calibration, decision curves, threshold-level classification, and SHapley Additive exPlanations (SHAP)-based explanation.
What is the implication, and what should change now?
• The model may support population-level early risk stratification, prognostic communication, and trial enrichment after external validation; it should not be used to infer treatment effects or directly assign treatment.
Introduction
Gastroesophageal and gastric malignancies remain major contributors to the global cancer burden. GLOBOCAN 2022 estimated nearly 20 million new cancer cases and 9.7 million cancer deaths worldwide, and cancers of the stomach and esophagus remained among the leading causes of cancer death (1). The global burden of esophageal adenocarcinoma has continued to rise in several populations, with projections suggesting further increases by 2040 (2). Population-based analyses have also reported persistent mortality from gastroesophageal junction and esophageal adenocarcinoma despite changes in incidence and treatment patterns (3).
Gastroesophageal junction adenocarcinoma (GEJA) occupies a clinically difficult boundary between distal esophageal and proximal gastric cancer. Surveillance, Epidemiology, and End Results (SEER)-based studies have characterized its clinicopathological patterns and survival, and a competing-risk model has been developed for esophagogastric junction adenocarcinoma (4,5). These studies are useful for general survival estimation, but they do not directly address a narrower decision-facing question at first metastatic presentation: which patients are most likely to die from cancer within the next few months.
Early mortality in metastatic GEJA is clinically important because it may reflect a combination of metastatic burden, aggressive histology, limited physiological reserve, treatment ineligibility, rapid clinical decline, and competing supportive-care needs. A six-month horizon is not intended to replace long-term survival prediction. It captures an early interval during which first-line systemic therapy is commonly initiated or withheld, adverse reserve becomes clinically apparent, and early palliative-care discussions may be particularly relevant. This endpoint therefore has value for early prognostic communication, trial enrichment, and stratified observational research, provided it is not misread as a treatment-decision rule.
Several nomograms have been proposed for early death or survival in stage IV or metastatic gastric cancer (6-8). Other models have focused on multi-organ metastases or liver-metastatic gastric cancer (9,10). These studies support the feasibility of registry-based early-death prediction, but they usually pool gastric subsites or metastatic patterns. A model specific to de novo metastatic GEJA may be more anatomically coherent, because GEJA differs from distal gastric cancer in epidemiology, treatment pathways, and clinical framing at presentation.
Treatment for gastroesophageal and gastric adenocarcinoma has evolved substantially. Perioperative chemotherapy, neoadjuvant chemoradiotherapy, and adjuvant nivolumab have reshaped care for localized or resectable disease (11-15). For advanced gastric or GEJA, trastuzumab establishes human epidermal growth factor receptor 2 (HER2)-directed therapy (16), immune checkpoint inhibitor combinations improve first-line treatment in selected or broad populations (17-20), trastuzumab deruxtecan changes later-line HER2-positive disease (21), CLDN18.2-directed therapy extends biomarker-guided treatment (22,23), ramucirumab-based therapy remains an important later-line option (24,25), and recent regulatory updates further emphasize the rapidly changing treatment context (26). These advances reinforce the need for early risk assessment, but registry variables cannot measure performance status, biomarker status, treatment line, dose intensity, or response in sufficient detail to support causal treatment inference.
Machine-learning models have recently been applied to gastric cancer metastasis and survival prediction (27,28). For clinical prediction, however, model development should not stop at apparent discrimination. Transparent reporting, calibration assessment, threshold-level classification, decision curve analysis, and interpretable variable-attribution methods are needed before a model can be considered reproducible or clinically meaningful (29-32). Because the best functional form was uncertain, this study compared model families with different assumptions, including conventional logistic regression, penalized regression, tree-based bagging, gradient boosting, and neural-network modelling. The objective was to develop and temporally validate an interpretable SEER-based prediction model for six-month cancer-specific mortality (CSM) in de novo metastatic GEJA using routinely available registry variables and organ-specific metastatic phenotypes. We present this article in accordance with the TRIPOD reporting checklist (available at https://jgo.amegroups.com/article/view/10.21037/jgo-2026-0513/rc).
Methods
Study design and data source
This was a population-based, retrospective, multivariable prediction-model development and temporal validation study. Data were obtained from the SEER research database using SEER*Stat case-listing export. SEER is a population-based cancer registry resource with standardized tumor and survival fields, but it remains a registry rather than a clinical-trial or hospital electronic health record dataset. Its sampling frame supports population-level analyses, while its limitations include restricted variable granularity, possible coding transitions across calendar years, unmeasured confounding, incomplete treatment detail, and potential selection or reporting bias. This study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments.
Study population and disease definition
Patients diagnosed from 2010 to 2021 were eligible. The case definition required C16.0-cardia, not otherwise specified (NOS) primary site, malignant behavior, positive histology, known age, and adenocarcinoma-related histology. The positive-histology requirement provided histopathologic confirmation. Autopsy-only and death-certificate-only reporting sources were excluded. To reduce ambiguity from multiple primaries, the cohort was restricted to a single primary or first primary malignancy according to SEER sequence-number fields.
De novo metastatic disease was defined using the best available metastasis evidence across SEER staging eras, including organ-specific metastasis fields, derived M-stage, extent-of-disease metastasis recode, and collaborative stage metastasis variables. Liver, lung, bone, and brain metastases were retained as separate predictors rather than collapsed into a metastatic count, because different metastatic sites may carry different implications for early mortality. Treatment variables were extracted from first-course treatment fields and used only as prognostic adjustment variables, not as causal treatment exposures.
Outcome definition and follow-up handling
The primary endpoint was CSM within six months after diagnosis. Cancer-specific death was determined from the SEER cause-specific death classification and survival months. Patients who survived beyond six months, died of non-cancer causes within six months, or were censored before the six-month endpoint were classified as event-free for six-month CSM. Patients without evaluable survival time or cause-of-death classification were excluded from the final analytic cohort. The endpoint should be interpreted with the usual caution for registry cause-of-death coding, because misclassification is possible in patients with comorbidities or competing terminal events.
Candidate predictors and preprocessing
Candidate predictors were selected a priori according to clinical relevance, SEER availability, cross-year definitional stability, and plausible availability at or near initial metastatic diagnosis. They included age at diagnosis, sex, race and ethnicity, marital status, year of diagnosis, histologic subtype, tumor grade, tumor size, liver metastasis, lung metastasis, bone metastasis, brain metastasis, primary-site surgery, radiotherapy, and chemotherapy. Variables with non-harmonizable definitions or insufficient usable information across the study period were not retained.
Tumor size was harmonized from collaborative stage tumor size for earlier diagnosis years and tumor size summary for later diagnosis years. Tumor grade was harmonized when available, while unknown grade was retained as a separate category. Missing proportions for candidate predictors are reported in Table 1. For model fitting, missing tumor size values were replaced with the development-cohort median and accompanied by an explicit tumor-size-missingness indicator. Unknown categories were retained when they reflected registry coding or clinically meaningful absence of information. Structural missingness in tumor grade, which was unavailable in the 2018-2021 temporal validation cohort because of SEER coding changes, was not imputed.
Table 1
| Characteristic | Overall (n=6,564) | Development (n=4,315) | Temporal validation (n=2,249) | P value |
|---|---|---|---|---|
| Age, years | 64.0 (56.0–72.0) | 64.0 (56.0–72.0) | 65.0 (57.0–73.0) | 0.009 |
| Tumor size, mm | 47.0 (30.0–61.0) | 46.0 (30.0–60.0) | 49.5 (30.0–63.0) | 0.19 |
| Sex | 0.28 | |||
| Male | 5,189 (79.1) | 3,428 (79.4) | 1,761 (78.3) | |
| Female | 1,375 (20.9) | 887 (20.6) | 488 (21.7) | |
| Race/ethnicity | <0.001 | |||
| Non-Hispanic White | 4,762 (72.5) | 3,206 (74.3) | 1,556 (69.2) | |
| Hispanic | 872 (13.3) | 535 (12.4) | 337 (15.0) | |
| Non-Hispanic Asian or Pacific Islander | 464 (7.1) | 284 (6.6) | 180 (8.0) | |
| Non-Hispanic Black | 407 (6.2) | 253 (5.9) | 154 (6.8) | |
| Other/unknown | 59 (0.9) | 37 (0.9) | 22 (1.0) | |
| Marital status | 0.001 | |||
| Married | 3,866 (58.9) | 2,571 (59.6) | 1,295 (57.6) | |
| Unmarried | 2,427 (37.0) | 1,544 (35.8) | 883 (39.3) | |
| Unknown | 271 (4.1) | 200 (4.6) | 71 (3.2) | |
| Histologic subtype | <0.001 | |||
| Adenocarcinoma, NOS | 5,276 (80.4) | 3,395 (78.7) | 1,881 (83.6) | |
| Signet ring cell carcinoma | 695 (10.6) | 516 (12.0) | 179 (8.0) | |
| Intestinal-type adenocarcinoma | 292 (4.4) | 191 (4.4) | 101 (4.5) | |
| Mucinous adenocarcinoma | 102 (1.6) | 79 (1.8) | 23 (1.0) | |
| Diffuse-type adenocarcinoma | 97 (1.5) | 55 (1.3) | 42 (1.9) | |
| Other adenocarcinoma | 79 (1.2) | 59 (1.4) | 20 (0.9) | |
| Tubular/papillary adenocarcinoma | 23 (0.4) | 20 (0.5) | 3 (0.1) | |
| Tumor grade | <0.001 | |||
| Unknown | 2,976 (45.3) | 727 (16.8) | 2,249 (100.0) | |
| Poorly differentiated | 2,337 (35.6) | 2,337 (54.2) | 0 (0.0) | |
| Moderately differentiated | 1,149 (17.5) | 1,149 (26.6) | 0 (0.0) | |
| Well differentiated | 102 (1.6) | 102 (2.4) | 0 (0.0) | |
| Tumor size missingness | 0.95 | |||
| Missing | 3,420 (52.1) | 2,247 (52.1) | 1,173 (52.2) | |
| Available | 3,144 (47.9) | 2,068 (47.9) | 1,076 (47.8) | |
| Bone metastasis | <0.001 | |||
| No | 5,429 (82.7) | 3,551 (82.3) | 1,878 (83.5) | |
| Yes | 947 (14.4) | 604 (14.0) | 343 (15.3) | |
| Unknown | 188 (2.9) | 160 (3.7) | 28 (1.2) | |
| Brain metastasis | <0.001 | |||
| No | 6,172 (94.0) | 4,011 (93.0) | 2,161 (96.1) | |
| Yes | 177 (2.7) | 126 (2.9) | 51 (2.3) | |
| Unknown | 215 (3.3) | 178 (4.1) | 37 (1.6) | |
| Liver metastasis | <0.001 | |||
| No | 3,177 (48.4) | 2,113 (49.0) | 1,064 (47.3) | |
| Yes | 3,249 (49.5) | 2,083 (48.3) | 1,166 (51.8) | |
| Unknown | 138 (2.1) | 119 (2.8) | 19 (0.8) | |
| Lung metastasis | <0.001 | |||
| No | 5,192 (79.1) | 3,351 (77.7) | 1,841 (81.9) | |
| Yes | 1,136 (17.3) | 767 (17.8) | 369 (16.4) | |
| Unknown | 236 (3.6) | 197 (4.6) | 39 (1.7) | |
| Primary-site surgery | 0.01 | |||
| No surgery | 5,978 (91.1) | 3,900 (90.4) | 2,078 (92.4) | |
| Primary-site surgery | 572 (8.7) | 403 (9.3) | 169 (7.5) | |
| Unknown | 14 (0.2) | 12 (0.3) | 2 (0.1) | |
| Radiotherapy | <0.001 | |||
| No/unknown radiation | 4,574 (69.7) | 2,941 (68.2) | 1,633 (72.6) | |
| Radiation | 1,986 (30.3) | 1,371 (31.8) | 615 (27.3) | |
| Unknown | 4 (0.1) | 3 (0.1) | 1 (0.0) | |
| Chemotherapy | 0.09 | |||
| Chemotherapy | 4,676 (71.2) | 3,044 (70.5) | 1,632 (72.6) | |
| No/unknown chemotherapy | 1,888 (28.8) | 1,271 (29.5) | 617 (27.4) | |
| 6-month cancer-specific death | <0.001 | |||
| No | 3,721 (56.7) | 2,374 (55.0) | 1,347 (59.9) | |
| Yes | 2,843 (43.3) | 1,941 (45.0) | 902 (40.1) | |
| 6-month other-cause death | 0.35 | |||
| No | 6,380 (97.2) | 4,200 (97.3) | 2,180 (96.9) | |
| Yes | 184 (2.8) | 115 (2.7) | 69 (3.1) |
Continuous variables are reported as median (interquartile range), and categorical variables are reported as n (%). P values compare development and temporal validation cohorts. Tumor grade was missing for all patients in the temporal validation cohort [2018–2021] due to changes in SEER data collection and cross-year coding availability. This limitation is addressed in the Discussion section. NOS, not otherwise specified; SEER, Surveillance, Epidemiology, and End Results.
Age and tumor size were represented flexibly using spline terms where appropriate. Categorical variables were encoded as indicator variables for model fitting. SEER chemotherapy is recorded as yes versus no/unknown; therefore, the variable was retained as its original registry treatment-status field and was not separated into true untreated and unknown categories.
Development, temporal validation, and model evaluation
Patients diagnosed in 2010–2017 formed the development cohort, and those diagnosed in 2018–2021 formed the temporal validation cohort. Within the development cohort, model fitting and Platt calibration were performed using an internal training-calibration split. The temporal validation cohort was held out from model fitting, calibration, and tuning, and was used only for final evaluation.
Representative model families were chosen to reflect different modelling assumptions: logistic regression as a conventional linear baseline, elastic net as a penalized regression method, random forest as a bagging-based tree ensemble, extreme gradient boosting (XGBoost) and gradient boosting machine as boosting-based tree ensembles, and multilayer perceptron as a neural-network classifier. Model selection prioritized the full validation profile, including temporal-validation discrimination, calibration, Brier score, decision curve analysis, threshold-level classification, and interpretability, rather than development-set area under the curve (AUC) alone.
AUC values with 95% confidence intervals (CIs) were estimated using DeLong’s method. Calibration was assessed using calibration intercept, calibration slope, Brier score, null Brier score, integrated calibration index, and calibration plots. Decision curve analysis was used to evaluate net benefit across clinically plausible threshold probabilities (31). SHapley Additive exPlanations (SHAP) analysis was used to explain the final tree-based model at the feature level (32).
No single AUC value was treated as sufficient to justify clinical use. In this oncology setting, temporal-validation AUC around 0.80 was considered compatible with population-level risk stratification only if calibration, threshold-level classification, and decision curves were acceptable. It was not considered sufficient for direct treatment assignment. Two prespecified risk cutoffs derived from tertiles of development-cohort predictions were applied unchanged to the temporal validation cohort: 26.5% for the low/intermediate boundary and 40.5% for the intermediate/high boundary. Sensitivity, specificity, positive predictive value, negative predictive value, and accuracy were calculated for these thresholds in temporal validation.
Sensitivity analyses
Prespecified sensitivity analyses evaluated a no-treatment model, the locked six-month risk score applied to 12-month CSM, candidate-model ranking, risk-group stability, and feature-contribution summaries. The no-treatment model removed primary-site surgery, radiotherapy, and chemotherapy to assess the model’s dependence on first-course treatment-status fields. Because SEER chemotherapy combines no and unknown chemotherapy, and because grade availability changed after 2018, both issues were addressed explicitly in interpretation rather than treated as causal findings.
Statistical analysis
Continuous variables are summarized as median (interquartile range) and were compared between the development and temporal validation cohorts using the Wilcoxon rank-sum test. Categorical variables are summarized as number (percentage) and were compared using Pearson's chi-square test; when expected cell counts were <5, Monte Carlo simulated chi-square P values based on 2,000 replicates were used. AUCs and 95% confidence intervals were estimated using DeLong’s method. All statistical tests were two-sided, and P<0.05 was considered statistically significant. Analyses were performed in R.
Results
Cohort derivation and baseline characteristics
The SEER*Stat export identified 22,772 patients with C16.0 cardia primary site, positive histology, malignant behavior, and diagnosis from 2010 to 2021. After restriction to adenocarcinoma-related histology, single or first primary malignancy, de novo metastatic disease, and evaluable six-month CSM status, 6,564 patients remained in the analytic cohort. The development cohort included 4,315 patients diagnosed in 2010–2017, and the temporal validation cohort included 2,249 patients diagnosed in 2018–2021 (Figure 1).
The median age was 64 years in the full cohort, and 5,189 patients (79.1%) were male. Non-Hispanic White patients accounted for 4,762 cases (72.5%). Adenocarcinoma, NOS was the dominant histologic subtype (5,276 patients, 80.4%), followed by signet ring cell carcinoma (695 patients, 10.6%). Liver metastasis was the most common organ-specific metastatic phenotype (3,249 patients, 49.5%), followed by lung metastasis (1,136 patients, 17.3%) and bone metastasis (947 patients, 14.4%). Chemotherapy was recorded in 4,676 patients (71.2%). Tumor grade was unavailable for all patients in the temporal validation cohort because of SEER coding availability after 2018; this was retained as a database limitation rather than imputed (Table 1).
Six-month CSM
Six-month cancer-specific death occurred in 1,941 of 4,315 patients (45.0%) in the development cohort and 902 of 2,249 patients (40.1%) in the temporal validation cohort. Six-month other-cause death occurred in 115 patients (2.7%) in development and 69 patients (3.1%) in temporal validation.
Candidate model comparison and final model selection
Model performance is summarized in Table 2. Logistic regression, elastic net, GBM, and XGBoost produced closely similar temporal-validation AUCs. Random forest showed apparent overfitting, with a development AUC of 0.924 but a temporal-validation AUC of 0.767 and a development calibration slope of 1.835, and was therefore not selected despite high apparent discrimination. MLP also showed a larger decline from development to temporal validation. The selected XGBoost model was retained because it combined competitive temporal-validation discrimination, acceptable calibration, decision-curve usefulness, threshold-level separation, and direct SHAP-based interpretability (Figures 2,3).
Table 2
| Model family | Model | Dataset | N | Events | Event rate, % | AUC (95% CI) | Brier score | Null Brier | Calibration intercept | Calibration slope | ICI | Log loss |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Logistic regression | Logistic regression | Development | 4,315 | 1,941 | 45 | 0.797 (0.784–0.810) | 0.179 | 0.247 | 0.177 | 1.03 | 0.03 | 0.534 |
| Logistic regression | Logistic regression | Temporal validation | 2,249 | 902 | 40.1 | 0.784 (0.764–0.803) | 0.174 | 0.24 | 0.21 | 1.036 | 0.039 | 0.528 |
| Elastic net | ElasticNet_alpha_0.75 | Development | 4,315 | 1,941 | 45 | 0.796 (0.783–0.809) | 0.18 | 0.247 | 0.168 | 1.011 | 0.032 | 0.536 |
| Elastic net | ElasticNet_alpha_0.75 | Temporal validation | 2,249 | 902 | 40.1 | 0.786 (0.767–0.806) | 0.173 | 0.24 | 0.201 | 1.026 | 0.031 | 0.524 |
| Random forest | RandomForest_mtry_6 | Development | 4,315 | 1,941 | 45 | 0.924 (0.916–0.932) | 0.126 | 0.247 | 0.563 | 1.835 | 0.119 | 0.405 |
| Random forest | RandomForest_mtry_6 | Temporal validation | 2,249 | 902 | 40.1 | 0.767 (0.746–0.787) | 0.179 | 0.24 | 0.085 | 1.061 | 0.047 | 0.541 |
| XGBoost | XGBoost_6 | Development | 4,315 | 1,941 | 45 | 0.808 (0.795–0.821) | 0.177 | 0.247 | 0.192 | 1.061 | 0.044 | 0.529 |
| XGBoost | XGBoost_6 | Temporal validation | 2,249 | 902 | 40.1 | 0.786 (0.767–0.806) | 0.173 | 0.24 | 0.18 | 1.13 | 0.026 | 0.524 |
| GBM | GBM_2 | Development | 4,315 | 1,941 | 45 | 0.813 (0.801–0.826) | 0.174 | 0.247 | 0.194 | 1.057 | 0.041 | 0.523 |
| GBM | GBM_2 | Temporal validation | 2,249 | 902 | 40.1 | 0.785 (0.765–0.804) | 0.173 | 0.24 | 0.112 | 1.019 | 0.026 | 0.524 |
| MLP | MLP_size5_decay0.1 | Development | 4,315 | 1,941 | 45 | 0.831 (0.819–0.843) | 0.168 | 0.247 | 0.297 | 1.366 | 0.057 | 0.51 |
| MLP | MLP_size5_decay0.1 | Temporal validation | 2,249 | 902 | 40.1 | 0.729 (0.708–0.751) | 0.203 | 0.24 | 0.089 | 0.827 | 0.067 | 0.6 |
AUC values with 95% CIs were estimated using DeLong’s method. Calibration metrics were calculated after Platt calibration where applicable. Random forest displayed strong optimism in the development cohort and was therefore not considered for final deployment despite its high apparent development AUC. AUC, area under the receiver operating characteristic curve; CI, confidence interval; GBM, gradient boosting machine; ICI, integrated calibration index; MLP, multilayer perceptron; XGBoost, extreme gradient boosting.
Final model performance, threshold behavior, and risk stratification
The selected XGBoost model achieved an AUC of 0.808 (95% CI, 0.795–0.821) in the development cohort and 0.786 (95% CI, 0.767–0.806) in the temporal validation cohort. In temporal validation, the Brier score was lower than the null Brier score (0.173 vs. 0.240), the calibration intercept was 0.180, the calibration slope was 1.130, and the integrated calibration index was 0.026 (Table 3).
Table 3
| Analysis | Dataset | Model | N | Events | Event rate, % | AUC (95% CI) | Brier score | Null Brier | Calibration intercept | Calibration slope | ICI |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Final selected model | Development | XGBoost_6 | 4,315 | 1,941 | 45 | 0.808 (0.795–0.821) | 0.177 | 0.247 | 0.192 | 1.061 | 0.044 |
| Final selected model | Temporal validation | XGBoost_6 | 2,249 | 902 | 40.1 | 0.786 (0.767–0.806) | 0.173 | 0.24 | 0.18 | 1.13 | 0.026 |
| No-treatment sensitivity model | Development | No-treatment sensitivity model | 4,315 | 1,941 | 45 | 0.683 (0.667–0.699) | 0.224 | 0.247 | 0.159 | 1.073 | 0.033 |
| No-treatment sensitivity model | Temporal validation | No-treatment sensitivity model | 2,249 | 902 | 40.1 | 0.650 (0.628–0.673) | 0.227 | 0.24 | 0.101 | 0.859 | 0.051 |
Treatment variables were used for prognostic adjustment, not causal inference. The no-treatment model intentionally removed primary-site surgery, radiotherapy, and chemotherapy variables to evaluate dependence on first-course treatment fields. AUC, area under the receiver operating characteristic curve; CI, confidence interval; ICI, integrated calibration index; XGBoost, extreme gradient boosting.
The threshold-specific classification profile is shown in Table 4. Using the 26.5% low/intermediate cutoff, classification of intermediate- or high-risk patients as test-positive yielded sensitivity of 76.5%, specificity of 62.1%, positive predictive value of 57.5%, and negative predictive value of 79.8%. Using the 40.5% intermediate/high cutoff, classification of high-risk patients as test-positive yielded lower sensitivity (56.7%) but higher specificity (90.2%) and positive predictive value (79.5%). This pattern supports use of the model for risk stratification rather than binary treatment assignment.
Table 4
| Threshold | Predicted risk cutoff | Positive classification | True positives | False positives | Sensitivity, % | Specificity, % | PPV, % | NPV, % | Accuracy, % |
|---|---|---|---|---|---|---|---|---|---|
| Low/intermediate threshold | ≥26.5% | Intermediate or high risk | 690 | 510 | 76.5 | 62.1 | 57.5 | 79.8 | 67.9 |
| Intermediate/high threshold | ≥40.5% | High risk | 511 | 132 | 56.7 | 90.2 | 79.5 | 75.7 | 76.7 |
Risk thresholds were derived from tertiles of final-model predictions in the development cohort and applied unchanged to temporal validation. NPV, negative predictive value; PPV, positive predictive value.
The locked risk groups separated observed six-month CSM in temporal validation. Observed event rates were 20.2% in the low-risk group, 32.1% in the intermediate-risk group, and 79.5% in the high-risk group. The mean predicted risks in these groups were 0.2007, 0.3101, and 0.7371, respectively. Calibration, clinical impact, and decision curve analyses are shown in Figure 4.
Sensitivity analyses and model explanation
Removing first-course treatment variables reduced temporal-validation discrimination from 0.786 to 0.650, indicating that treatment-status fields carried substantial prognostic information but also confirming that the final model relied partly on registry treatment variables (Table 3). The 12-month sensitivity analysis showed that the locked six-month risk score remained associated with longer-horizon CSM but performed less strongly for 12-month prediction, with a temporal-validation AUC of 0.732. This supported the intended use of the model as an early six-month risk tool rather than a general long-term survival model.
SHAP analysis identified chemotherapy status, primary-site surgery status, liver metastasis, tumor grade, age, bone metastasis, diagnosis year, lung metastasis, signet ring cell histology, marital status, radiotherapy, tumor size, and brain metastasis as influential predictors (Figure 5). The large contributions of chemotherapy and surgery fields should be interpreted as prognostic information captured in registry treatment-status variables, not as evidence of causal treatment benefit or harm.
Discussion
Principal findings
In this population-based study, an interpretable machine-learning model was developed and temporally validated to estimate six-month CSM in patients with de novo metastatic GEJA. The model achieved moderate temporal-validation discrimination, acceptable calibration, a Brier score lower than the null Brier score, and clearly separated risk strata. The high-risk group had an observed six-month CSM rate of nearly 80%, whereas the low-risk group had an observed rate of approximately 20%. These results suggest that registry variables can stratify early CSM risk at the population level when model development is accompanied by chronological validation, calibration assessment, threshold-level evaluation, and cautious interpretation.
Comparison with previous models
The study addresses a distinct clinical gap. Prior SEER-based work in esophagogastric junction adenocarcinoma has focused largely on overall survival, competing-risk survival, or broad clinicopathologic description (4,5). Early-death models in metastatic gastric cancer have shown that registry variables can identify high-risk patients, but those studies generally pooled gastric cancer subsites or focused on specific metastatic patterns such as liver metastasis or multi-organ spread (6-10). By restricting the cohort to C16.0 GEJA, requiring de novo metastatic disease, and modelling liver, lung, bone, and brain metastases separately, the present study avoids a generic metastatic-count approach and remains closer to the clinical anatomy of GEJA.
The model-selection results also show why algorithm ranking by apparent AUC alone is inadequate. Random forest had a high development AUC but weaker temporal validation and marked calibration optimism, a pattern compatible with overfitting. XGBoost was selected because it balanced discrimination, calibration, Brier score, decision-curve performance, threshold behavior, and interpretability. The similar temporal-validation AUCs of elastic net, GBM, and XGBoost indicate that the endpoint signal was not confined to one algorithm; the final selection was based on the overall validation profile rather than a marginal AUC difference.
Clinical interpretation
The six-month horizon should be interpreted as an early prognosis window, not as a surrogate for treatment effect. In metastatic GEJA, early death may arise from extensive metastatic burden, aggressive histology, comorbidity, frailty, rapidly progressive symptoms, limited access to systemic therapy, or inability to tolerate treatment. The model can identify groups with materially different early mortality risk, which may be useful for risk communication, observational stratification, and trial enrichment. However, it should not be used to deny or assign systemic therapy, surgery, radiotherapy, or palliative-care interventions.
The threshold analysis provides a more clinically interpretable view than AUC alone. The low/intermediate cutoff favored sensitivity and may be more useful for identifying a broad group requiring closer early assessment. The intermediate/high cutoff favored specificity and positive predictive value, identifying a smaller subgroup with very high observed early CSM. These trade-offs reinforce that model use depends on the clinical purpose. A threshold that is useful for screening or research enrichment is not necessarily appropriate for high-stakes treatment decisions.
The SHAP analysis supported several clinically plausible patterns, including contributions from organ-specific metastases, age, histology, grade, and registry treatment-status fields. The treatment-related findings require particular restraint. SEER chemotherapy is coded as yes versus no/unknown. The no/unknown group may include patients who were too ill for chemotherapy, declined treatment, received treatment outside reporting pathways, died rapidly before treatment, or had incomplete documentation. Similarly, primary-site surgery status may reflect selection and documentation processes. These variables were retained for prognostic adjustment but cannot be interpreted causally.
Limitations
The most important limitation is the absence of validation outside SEER. Temporal validation is stricter than a random split because it tests performance across calendar time and coding transitions, but it remains internal to the same registry system, with the same broad data architecture, variable definitions, and reporting practices. It does not establish transportability to other registries, countries, hospital networks, or treatment eras. Independent validation in non-SEER datasets and local recalibration would be required before clinical workflow use.
Several registry limitations also affect interpretation. SEER lacks Eastern Cooperative Oncology Group performance status, laboratory values, nutritional status, liver and renal function, inflammatory markers, detailed tumor burden, HER2 status, programmed death-ligand 1 expression, microsatellite instability, CLDN18.2 expression, treatment line, chemotherapy regimen, immunotherapy exposure, dose intensity, response, and supportive-care details. Cause-specific death classification may be imperfect, especially in patients with comorbidities or competing terminal events. Tumor grade was unavailable for all temporal-validation patients because of SEER coding changes after 2018. This missingness was treated as structural database missingness rather than imputed.
The study should therefore be understood as a population-based early-risk stratification analysis, not as a bedside treatment-assignment instrument. A future model that combines SEER-like population coverage with prospective clinical variables, biomarker data, treatment sequence, and external validation could improve calibration, clinical interpretability, and transportability.
Conclusions
A SEER-based XGBoost model was developed and temporally validated to estimate six-month CSM in de novo metastatic GEJA. The model provided moderate discrimination, acceptable calibration, clinically distinct risk strata, and interpretable feature attribution. Its appropriate role is early prognostic risk stratification and research support after external validation and local recalibration, not causal treatment inference or direct treatment assignment.
Acknowledgments
The authors thank the SEER Program registries for providing access to population-based cancer data.
Footnote
Reporting Checklist: The authors have completed the TRIPOD reporting checklist. Available at https://jgo.amegroups.com/article/view/10.21037/jgo-2026-0513/rc
Peer Review File: Available at https://jgo.amegroups.com/article/view/10.21037/jgo-2026-0513/prf
Funding: This study was supported by
Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://jgo.amegroups.com/article/view/10.21037/jgo-2026-0513/coif). The authors have no conflicts of interest to declare.
Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. This study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments.
Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.
References
- Bray F, Laversanne M, Sung H, et al. Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA Cancer J Clin 2024;74:229-63. [Crossref] [PubMed]
- Morgan E, Soerjomataram I, Rumgay H, et al. The Global Landscape of Esophageal Squamous Cell Carcinoma and Esophageal Adenocarcinoma Incidence and Mortality in 2020 and Projections to 2040: New Estimates From GLOBOCAN 2020. Gastroenterology 2022;163:649-658.e2. [Crossref] [PubMed]
- Agarwal S, Bell MG, Dhaliwal L, et al. Population Based Time Trends in the Epidemiology and Mortality of Gastroesophageal Junction and Esophageal Adenocarcinoma. Dig Dis Sci 2024;69:246-53. [Crossref] [PubMed]
- Liu X, Jiang Q, Yue C, et al. Clinicopathological Characteristics and Survival Predictions for Adenocarcinoma of the Esophagogastric Junction: A SEER Population-Based Retrospective Study. Int J Gen Med 2021;14:10303-14. [Crossref] [PubMed]
- Wang T, Wu Y, Zhou H, et al. Development and validation of a novel competing risk model for predicting survival of esophagogastric junction adenocarcinoma: a SEER population-based study and external validation. BMC Gastroenterol 2021;21:38. [Crossref] [PubMed]
- Zhu Y, Fang X, Wang L, et al. A Predictive Nomogram for Early Death of Metastatic Gastric Cancer: A Retrospective Study in the SEER Database and China. J Cancer 2020;11:5527-35. [Crossref] [PubMed]
- Feng Y, Guo K, Jin H, et al. A Predictive Nomogram for Early Mortality in Stage IV Gastric Cancer. Med Sci Monit 2020;26:e923931. [Crossref] [PubMed]
- Yang Y, Chen ZJ, Yan S. The incidence, risk factors and predictive nomograms for early death among patients with stage IV gastric cancer: a population-based study. J Gastrointest Oncol 2020;11:964-82. [Crossref] [PubMed]
- Wang T, Liu C, Wang W, et al. Development and validation of a Surveillance, Epidemiology, and End Results (SEER)-based prognostic nomogram for predicting survival in gastric cancer with multi-organ metastases. Transl Cancer Res 2022;11:1534-51. [Crossref] [PubMed]
- Meng N, Niu X, Wu J, et al. Development and validation of nomogram models for predicting overall survival and cancer-specific survival in gastric cancer patients with liver metastases: a cohort study based on the SEER database. Am J Cancer Res 2024;14:2272-86. [Crossref] [PubMed]
- Al-Batran SE, Homann N, Pauligk C, et al. Perioperative chemotherapy with fluorouracil plus leucovorin, oxaliplatin, and docetaxel versus fluorouracil or capecitabine plus cisplatin and epirubicin for locally advanced, resectable gastric or gastro-oesophageal junction adenocarcinoma (FLOT4): a randomised, phase 2/3 trial. Lancet 2019;393:1948-57. [Crossref] [PubMed]
- Cunningham D, Allum WH, Stenning SP, et al. Perioperative chemotherapy versus surgery alone for resectable gastroesophageal cancer. N Engl J Med 2006;355:11-20. [Crossref] [PubMed]
- van Hagen P, Hulshof MC, van Lanschot JJ, et al. Preoperative chemoradiotherapy for esophageal or junctional cancer. N Engl J Med 2012;366:2074-84. [Crossref] [PubMed]
- Shapiro J, van Lanschot JJB, Hulshof MCCM, et al. Neoadjuvant chemoradiotherapy plus surgery versus surgery alone for oesophageal or junctional cancer (CROSS): long-term results of a randomised controlled trial. Lancet Oncol 2015;16:1090-8. [Crossref] [PubMed]
- Kelly RJ, Ajani JA, Kuzdzal J, et al. Adjuvant Nivolumab in Resected Esophageal or Gastroesophageal Junction Cancer. N Engl J Med 2021;384:1191-203. [Crossref] [PubMed]
- Bang YJ, Van Cutsem E, Feyereislova A, et al. Trastuzumab in combination with chemotherapy versus chemotherapy alone for treatment of HER2-positive advanced gastric or gastro-oesophageal junction cancer (ToGA): a phase 3, open-label, randomised controlled trial. Lancet 2010;376:687-97. [Crossref] [PubMed]
- Janjigian YY, Shitara K, Moehler M, et al. First-line nivolumab plus chemotherapy versus chemotherapy alone for advanced gastric, gastro-oesophageal junction, and oesophageal adenocarcinoma (CheckMate 649): a randomised, open-label, phase 3 trial. Lancet 2021;398:27-40. [Crossref] [PubMed]
- Kang YK, Chen LT, Ryu MH, et al. Nivolumab plus chemotherapy versus placebo plus chemotherapy in patients with HER2-negative, untreated, unresectable advanced or recurrent gastric or gastro-oesophageal junction cancer (ATTRACTION-4): a randomised, multicentre, double-blind, placebo-controlled, phase 3 trial. Lancet Oncol 2022;23:234-47. [Crossref] [PubMed]
- Rha SY, Oh DY, Yañez P, et al. Pembrolizumab plus chemotherapy versus placebo plus chemotherapy for HER2-negative advanced gastric cancer (KEYNOTE-859): a multicentre, randomised, double-blind, phase 3 trial. Lancet Oncol 2023;24:1181-95. [Crossref] [PubMed]
- Janjigian YY, Kawazoe A, Bai Y, et al. Pembrolizumab plus trastuzumab and chemotherapy for HER2-positive gastric or gastro-oesophageal junction adenocarcinoma: interim analyses from the phase 3 KEYNOTE-811 randomised placebo-controlled trial. Lancet 2023;402:2197-208. [Crossref] [PubMed]
- Shitara K, Bang YJ, Iwasa S, et al. Trastuzumab Deruxtecan in Previously Treated HER2-Positive Gastric Cancer. N Engl J Med 2020;382:2419-30. [Crossref] [PubMed]
- Shitara K, Lordick F, Bang YJ, et al. Zolbetuximab plus mFOLFOX6 in patients with CLDN18.2-positive, HER2-negative, untreated, locally advanced unresectable or metastatic gastric or gastro-oesophageal junction adenocarcinoma (SPOTLIGHT): a multicentre, randomised, double-blind, phase 3 trial. Lancet 2023;401:1655-68. [Crossref] [PubMed]
- Shah MA, Shitara K, Ajani JA, et al. Zolbetuximab plus CAPOX in CLDN18.2-positive gastric or gastroesophageal junction adenocarcinoma: the randomized, phase 3 GLOW trial. Nat Med 2023;29:2133-41. [Crossref] [PubMed]
- Fuchs CS, Tomasek J, Yong CJ, et al. Ramucirumab monotherapy for previously treated advanced gastric or gastro-oesophageal junction adenocarcinoma (REGARD): an international, randomised, multicentre, placebo-controlled, phase 3 trial. Lancet 2014;383:31-9. [Crossref] [PubMed]
- Wilke H, Muro K, Van Cutsem E, et al. Ramucirumab plus paclitaxel versus placebo plus paclitaxel in patients with previously treated advanced gastric or gastro-oesophageal junction adenocarcinoma (RAINBOW): a double-blind, randomised phase 3 trial. Lancet Oncol 2014;15:1224-35. [Crossref] [PubMed]
- U.S. Food and Drug Administration. FDA approves pembrolizumab for HER2 positive gastric or gastroesophageal junction adenocarcinoma expressing PD-L1 (CPS >=1). Published March 19, 2025. Accessed June 30, 2026. Available online: https://www.fda.gov/drugs/resources-information-approved-drugs/fda-approves-pembrolizumab-her2-positive-gastric-or-gastroesophageal-junction-adenocarcinoma
- Qin X, Qiu B, Ge L, et al. Applying machine learning techniques to predict the risk of distant metastasis from gastric cancer: a real world retrospective study. Front Oncol 2024;14:1455914. [Crossref] [PubMed]
- Zhang C, Zhang Y, Yang YH, et al. Machine learning models for predicting one-year survival in patients with metastatic gastric cancer who experienced upfront radical gastrectomy. Front Mol Biosci 2022;9:937242. [Crossref] [PubMed]
- Collins GS, Reitsma JB, Altman DG, et al. Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (TRIPOD): the TRIPOD Statement. BMC Med 2015;13:1. [Crossref] [PubMed]
- Collins GS, Moons KGM, Dhiman P, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ 2024;385:e078378. [Crossref] [PubMed]
- Vickers AJ, Elkin EB. Decision curve analysis: a novel method for evaluating prediction models. Med Decis Making 2006;26:565-74. [Crossref] [PubMed]
- Lundberg SM, Erion G, Chen H, et al. From Local Explanations to Global Understanding with Explainable AI for Trees. Nat Mach Intell 2020;2:56-67. [Crossref] [PubMed]

