Radiomic prediction of early progression at 2 years post-treatment in patients with resectable rectal cancer based on rectal tumor and mesentery characteristics: a two-center study
Highlight box
Key findings
• We constructed a deep radiomics model that integrates preoperative magnetic resonance imaging (MRI) radiomics features and deep learning features to predict early progression (EP) with resectable rectal cancer. This study adopted multi-cohort verification based on training, internal test, and external validation datasets. SHapley Additive exPlanations (SHAP) analysis was performed to visualize and interpret the optimal model, further improving its clinical practicability and acceptability.
What is known and what is new?
• Current routine clinical evaluation cannot accurately stratify rectal cancer patients at high risk of EP within 2 years after surgery.
• We innovatively combined imaging features of both rectal tumor and mesorectal tissue to build a deep radiomics model. The proposed model effectively identifies patients at high risk of EP, and SHAP visualization enhances model interpretability, achieving better predictive performance than conventional clinical assessment.
What is the implication, and what should change now?
• This non-invasive MRI-based deep radiomics model enables individualized risk stratification and postoperative surveillance for rectal cancer patients. Integrating tumor and mesorectal signature analysis can further optimize preoperative risk evaluation in clinical practice.
Introduction
Colorectal cancer (CRC) is currently one of the most common types of cancer worldwide. According to GLOBOCAN 2022 data, CRC ranks third globally in incidence and second in mortality (1). Rectal cancer accounts for about one-third of all CRCs (2). The primary surgical approach for rectal cancer is total mesorectal excision. However, the incidence of postoperative recurrence or metastasis remains high (3), with about 50% of high-risk patients still facing recurrence or distant metastasis. The British physician Bill Heald first introduced the anatomical concept of the rectal mesentery. Since then, new insights into its structure and function have emerged (4). The rectal mesentery contains a rich network of blood vessels, lymphatic vessels, nerves, and lymphatic tissue. These structures facilitate tumor metastasis and spread (5), making the mesentery a major factor affecting rectal cancer prognosis. The traditional tumor-node-metastasis (TNM) staging system remains the gold standard for assessing clinical prognosis. However, it does not account for tumor heterogeneity, so its accuracy is suboptimal. There is an urgent need for a more accurate and thorough assessment system. If patient recurrence and metastasis could be non-invasively predicted using imaging data of rectal tumors and the rectal mesentery before treatment, prognosis and quality of life could be greatly improved.
With advances in modern technology, artificial intelligence (AI) has been integrated into the field of medical image analysis, with radiomics and deep learning serving as the two mainstream technical approaches for AI-based quantitative image analysis. Radiomics is a quantitative analysis technique based on conventional medical images that extracts a large number of textural features from image data and correlates them with clinical issues to improve diagnostic and predictive performance. Deep learning, on the other hand, relies on convolutional neural networks to autonomously learn deep image features. It can adaptively match tumors, target the extraction of lesion-specific features, and accommodate the assessment needs of tumors that undergo deformation. The use of AI for automated quantitative analysis of medical images can effectively reduce inter-rater variability, improve the stability and reproducibility of assessments, and provide a viable solution for the standardized quantitative imaging assessment of gastrointestinal tumors (6). Both radiomics and deep learning have produced important findings in the assessment of rectal cancer staging, genetic status, treatment response, lymph node metastasis, and prognosis (7-11). Existing studies have largely focused on extracting imaging features from the primary tumor, but have paid little attention to the rectal mesentery, a key anatomical region for prognosis. The rectal mesentery serves as the primary site for occult microinvasion, extramural vascular invasion, and microscopic mesenteric metastases. Pathological changes such as mesenteric adipose fibrosis, disruption of the inflammatory microenvironment, and vascular invasion directly increase the risk of lymph node metastasis and postoperative local recurrence, thereby possessing independent prognostic value (12-14). These mesenchymal characteristics can indirectly reflect the primary tumor’s invasive potential and the burden of occult micrometastases, thereby helping to predict a patient’s risk of postoperative early progression (EP) (15,16). Therefore, this study simultaneously incorporates imaging features from both the primary tumor and the rectal mesentery, with the aim of addressing the limitation that intratumoral features alone cannot fully characterize the tumor invasion microenvironment. International and domestic guidelines for CRC diagnosis and treatment recommend high-resolution rectal magnetic resonance imaging (MRI) as the preferred imaging modality (17). High-resolution MRI provides exceptional soft-tissue resolution. It clearly depicts the tumor’s location, size, and depth of invasion. MRI enables precise assessment of circumferential margins and supplies critical evidence for clinical staging. Traditional imaging relies mainly on the radiologist’s visual inspection and subjective judgment. As a result, it cannot reveal in-depth information about the lesion’s internal pathology (the biological changes within the tumor) and immunological characteristics (how the tumor interacts with the body’s immune system). Bioinformatics and deep learning provide non-invasive, accurate, in-depth, and multidimensional assessment of rectal cancer. Yet, AI faces clinical challenges, especially regarding interpretability. Interpretability—the ability to understand how and why an AI makes certain decisions—is important for clinicians to grasp the rationale behind AI-assisted diagnostic and treatment recommendations (18). In the future, such preoperative imaging assessment models could be further expanded to cover additional application scenarios. Drawing on the development approach used for real-time imaging navigation in minimally invasive surgery, these models could gradually evolve from preoperative risk stratification to integrated decision support for intraoperative minimally invasive procedures, thereby advancing imaging AI technology from retrospective research to real-time clinical implementation during surgery (19).
This study mainly uses radiomics and deep learning to analyze rectal tumors and the rectal mesentery, with the goal of developing a model to predict EP in rectal cancer patients 2 years after treatment. To enhance acceptability and usefulness, the model is visualized and interpreted using Shapley Additive exPlanations (SHAP) techniques. Ultimately, this approach aims to provide a non-invasive adjunctive assessment tool for rectal cancer patients before surgery. However, traditional prognostic assessment indicators (such as disease-free survival and overall survival) primarily focus on long-term survival rates and rely on Cox regression analysis; they are unable to identify high-risk patients at an early stage. Approximately 60–80% of rectal cancer recurrences occur within 2 years after surgery (20), therefore, this study selected “EP within 2 years” as a binary composite endpoint. By identifying patients at higher risk of EP, this model could help clinicians personalize treatment strategies, such as intensifying follow-up, considering additional therapies, or tailoring surgical and adjuvant treatment plans. Integration of the model into routine clinical workflows has the potential to support more informed decision-making and improve individualized patient management. We present this article in accordance with the TRIPOD reporting checklist (available at https://jgo.amegroups.com/article/view/10.21037/jgo-2026-0451/rc).
Methods
Study population
The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. The study was approved by the Ethics Committees of The First Affiliated Hospital of Soochow University [approval No. (2025) Ethical Review No. 739] and the Suzhou Hospital Affiliated to Nanjing Medical University (approval No. K-2025-206-K01). Because this was a retrospective study, patients did not need to sign informed consent forms. The study included 255 patients with rectal cancer from The First Affiliated Hospital of Soochow University (Center 1) (January 2018–September 2022) and 69 patients from the Suzhou Hospital Affiliated to Nanjing Medical University (Center 2) (January 2018–September 2023). Data from Center 1 were used for model training and internal testing. Data from Center 2 served as the external validation set. In Center 1, 68 patients experienced an EP, and 187 did not; in Center 2, 13 and 56, respectively. Inclusion criteria: (I) pathologically confirmed rectal cancer; (II) complete surgical records for rectal cancer; (III) MRI images of sufficient quality for analysis; and (IV) complete clinical, pathological, and follow-up data. The follow-up must be at least 2 years for those without an endpoint event. Exclusion criteria: (I) history of other or concurrent malignancies at diagnosis; (II) prior antitumor therapy (such as radiotherapy or chemotherapy) before MRI; and (III) poor image quality or severe artifacts that interfere with evaluation.
Selection of clinical data
Clinical data included patient age, sex, comorbidities, pathology, family history, smoking history, alcohol consumption history, carcinoembryonic antigen (CEA), carbohydrate antigen 19-9 (CA19-9), start date of neoadjuvant therapy, date of surgery, and date of MRI scan. In the training set, 27 patients had missing records for CEA and CA19-9 (with a missing rate of 10%). Multiple imputation was used to handle the missing data, with the outcome variable included in the imputation model. A total of 10 imputed datasets were generated, and the final results were pooled using Rubin’s rules. Strict controls ensured that the interval between the MRI and the start of treatment did not exceed 14 days. Tumor recurrence or metastasis was assessed by reviewing post-treatment imaging, endoscopic, and biopsy records. The date of the first neoadjuvant therapy or surgery was set as the starting point. Tumor recurrence or metastasis within 2 years of treatment was defined as EP; otherwise, as non-EP. In this study, patients in the EP group were followed until an endpoint event. Those in the non-EP group were followed for more than 24 months.
MRI image acquisition
This study used MRI systems from Philips, Siemens, and GE, as well as phased-array coils (devices for improving image quality). The images acquired were T2-weighted spin-echo images in the transverse plane without fat suppression (a mode that retains fat signal for anatomical clarity). Detailed scanning parameters are shown in Table 1.
Table 1
| Manufacturer | Type | Sequence | TR (ms) | TE (ms) | Field strength (T) | Thickness (mm) | Matrix |
|---|---|---|---|---|---|---|---|
| Philips | Ingenia Ambition | SE | 2,000 | 90 | 1.5 | 3–4 | 180×180 |
| Ingenia | SE | 4,000 | 80 | 3 | 3–4 | 220×200 | |
| Siemens | Skyra | SE | 3,500 | 85 | 3 | 3–4 | 240×320 |
| Verio | SE | 3,400 | 92 | 3 | 3–4 | 380×300 | |
| GE | HD X | SE | 2,180 | 110 | 3 | 3–4 | 256×320 |
MRI, magnetic resonance imaging; SE, spin echo; TE, echo time; TR, repetition time.
Image segmentation
Using ITK-SNAP software (version 4.2.0; https://www.itk-snap.org/), three-dimensional (3D) volumes of interest (VOIs) for the tumor and mesentery were delineated layer by layer on T2-weighted images (T2WI), carefully distinguishing between the tumor, the rectal mesentery, and surrounding tissues. All VOI delineations were performed by Attending Physician A, who has 6 years of experience in MRI diagnosis. To assess the consistency of the delineations, 30 patients were randomly selected and independently delineated by another associate chief physician, Physician B, who had 10 years of experience. One month later, Physician A repeated the delineation for these 30 cases. The intra-class correlation coefficient (ICC) was used to assess intra-observer and inter-observer agreement, respectively; an ICC between 0.60 and 0.79 was considered to indicate good agreement. The intra-observer ICC was 0.738 [95% confidence interval (CI): 0.726–0.751], while the inter-observer ICC was 0.727 (95% CI: 0.714–0.740), indicating good consistency and reproducibility. Two physicians delineated the regions of interest (ROIs) without access to the patients’ clinical or pathological information; the delineated ROIs are shown in Figure 1.
Feature extraction
First, the images undergo preprocessing. To minimize the effects of the machine model, scanning parameters, and field strength, features are extracted only after preprocessing. An N4 bias field correction is applied to correct for signal inhomogeneity in the MRI data. Image voxels are resampled to a uniform 1×1×1 mm3 resolution to ensure consistent scale. Finally, the data undergoes Z-score normalization to convert it to a distribution that approximates a standard normal distribution.
In this study, patients from Center 1 were randomly divided into a training set (n=204) and a test set (n=51) in an 8:2 ratio, while patients from Center 2 served as the external validation set (n=69). For feature extraction, we first used the PyRadiomics package (version 3.1.0) to extract radiomic features. For deep learning features, the model employs a 2.5-dimensional (2.5D) multi-view input strategy, which involves extracting lesion images from the maximum sections in the axial, sagittal, and coronal planes and fusing them into a three-channel input for the network. The training parameters are set to batch_size =64 and epochs =100. Model training was carried out using four network architectures: ResNet-50, DenseNet-121, ResNet-18, and CrossFormer. Based on performance evaluations of each architecture, the optimal model was selected, and deep learning features were extracted from its average pooling layer.
Feature selection and model development
Based on the results of the intra-observer and inter-observer ICC tests, features with poor reproducibility (ICC <0.7) were excluded in advance; the remaining imageomics features and deep learning features were first standardized using Z-scores. Subsequently, features with non-significant intergroup differences were excluded via statistical testing (features with P<0.05 were retained). Correlations between features were calculated; if the correlation coefficient between two features was >0.8, only one of them was retained. Next, we performed five-fold cross-validation on the training set and used the least absolute shrinkage and selection operator (LASSO) regression to further identify the subset of features most valuable for prediction. Before training the model, the synthetic minority oversampling technique (SMOTE) was used to oversample samples from the minority class, thereby addressing the imbalance between the number of positive and negative samples. The ratio of negative to positive samples was set to 1:1 to resolve the imbalance between the number of positive and negative samples. This study evaluated eight machine learning classifiers, including: support vector machines (SVMs), K-nearest neighbors (KNNs), random forests (RFs), extremely randomized trees (ExtraTrees), extreme gradient boosting (XGBoost), light gradient boosting machine (LightGBM), multi-layer perceptron (MLP), and logistic regression (LR). Models were built using Python, and the best model was selected by comparing the area under the receiver operating characteristic (ROC) curve (AUC) for each model.
Modeling clinical characteristics
Univariate LR was used to select variables, followed by the construction of clinical prediction models using the eight machine learning classifiers described above. Multivariate LR was then used to identify independent predictors.
Statistical analysis
Statistical analysis was performed using SPSS 25.0 software. Continuous variables were assessed for normality using the Kolmogorov-Smirnov test. Continuous variables that followed a normal distribution were reported as the mean ± standard deviation, and differences between groups were compared using the independent-samples t-test. Continuous variables with skewed distributions were reported as the median (interquartile range), and differences between groups were compared using the Mann-Whitney U test. Categorical data were expressed as counts (percentages), and differences between groups were compared using the Chi-squared test or Fisher’s exact test. A two-tailed P value <0.05 was considered statistically significant. ROC curves were plotted, and AUC values were calculated.
Model explainability analysis
This study uses the SHAP method to conduct interpretability analysis of machine learning models, intending to quantify each feature’s contribution to the model’s predictions. By calculating SHAP scores for each sample, the model’s predicted values are decomposed into the additive contributions of individual features. These results are visualized at both the global (feature significance ranking) and local (single-sample decision path) levels. Specifically, global attribute importance is presented in a summary plot, while the decision logic for individual samples is illustrated using waterfall plots and force maps to reveal the factors driving each sample’s predictions.
Results
Baseline data
Among the 324 patients with rectal cancer, 81 had EP and 243 did not. Clinical baseline data showed no statistically significant differences between the EP groups in terms of age, CA19-9, CEA, gender, history of diabetes, history of hypertension, family history, smoking history, or whether patients had received neoadjuvant therapy and were treated for resectable rectal cancer 2 years post-treatment (P>0.05). The baseline data for the rectal cancer patients included in this study are shown in Table 2.
Table 2
| Clinical variables | Training cohort | Test cohort | External validation cohort | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| No EP (n=149) | EP (n=55) | P | No EP (n=38) | EP (n=13) | P | No EP (n=56) | EP (n=13) | P | |||
| Age (years) | 61.37±10.59 | 62.05±13.04 | 0.73 | 60.29±10.12 | 59.23±15.25 | 0.82 | 64.88±8.84 | 65.23±6.70 | 0.88 | ||
| CA19-9 (U/mL) | 0.80 | 0.31† | 0.02† | ||||||||
| ≤37 | 132 (88.59) | 48 (87.27) | 33 (86.64) | 13 (100.00) | 55 (98.21) | 10 (76.92) | |||||
| >37 | 17 (11.41) | 7 (12.73) | 5 (13.16) | 0 | 1 (1.79) | 3 (23.08) | |||||
| CEA | 0.94 | 0.91 | 0.09 | ||||||||
| 0–5 (non-smokers), 0–10 (smokers) | 113 (75.84) | 42 (76.36) | 26 (68.42) | 8 (61.54) | 42 (75.00) | 6 (46.15) | |||||
| >5 (non-smokers), >10 (smokers) | 36 (24.16) | 13 (23.64) | 12 (31.58) | 5 (38.46) | 14 (25.00) | 7 (53.85) | |||||
| Gender | 0.64 | 0.056 | 0.01† | ||||||||
| Female | 49 (32.89) | 20 (36.36) | 12 (31.58) | 8 (61.54) | 19 (33.93) | 0 | |||||
| Male | 100 (67.11) | 35 (63.64) | 26 (68.42) | 5 (38.46) | 37 (66.07) | 13 (100.00) | |||||
| DM | 0.89 | 0.45† | >0.99 | ||||||||
| No | 141 (94.63) | 53 (96.36) | 37 (100.00) | 12 (92.31) | 48 (85.71) | 11 (84.62) | |||||
| Yes | 8 (5.37) | 2 (3.64) | 1 (2.63) | 1 (7.69) | 8 (14.29) | 2 (15.38) | |||||
| HTN | 0.58 | 0.17 | 0.39 | ||||||||
| No | 94 (63.09) | 37 (67.27) | 21 (55.26) | 10 (76.92) | 27 (48.21) | 8 (61.54) | |||||
| Yes | 55 (36.91) | 18 (32.73) | 17 (44.74) | 3 (23.08) | 29 (51.79) | 5 (38.46) | |||||
| Family history | 0.91 | >0.99† | >0.99† | ||||||||
| No | 144 (96.64) | 54 (98.18) | 38 (100.00) | 13 (100.00) | 53 (94.64) | 13 (100.00) | |||||
| Yes | 5 (3.36) | 1 (1.82) | 0 | 0 | 3 (5.36) | 0 | |||||
| Smoking history | 0.69 | 0.02† | >0.99† | ||||||||
| Non-smoker | 128 (85.91) | 46 (83.64) | 26 (68.42) | 13 (100.00) | 52 (92.86) | 13 (100.00) | |||||
| Smoker | 21 (14.09) | 9 (16.36) | 12 (31.58) | 0 | 4 (7.14) | 0 | |||||
| NAT | 0.55 | >0.99 | >0.99† | ||||||||
| No | 125 (83.89) | 48 (87.27) | 33 (86.84) | 11 (84.62) | 53 (94.64) | 13 (100.00) | |||||
| Yes | 24 (16.11) | 7 (12.73) | 5 (13.16) | 2 (15.38) | 3 (5.36) | 0 | |||||
Data are presented as mean ± standard deviation or n (%). †, the Fisher test was used. CA19-9, carbohydrate antigen 19-9; CEA, carcinoembryonic antigen; DM, diabetes mellitus; EP, early progression; HTN, hypertension; NAT, neoadjuvant therapy.
Selection of radiomic and deep learning features
A total of 2,981 features were extracted to construct the radiomics model. After screening using t-tests or U-tests, the 9 most valuable radiomics features were ultimately retained. Regarding deep learning features, based on a comparative evaluation of network architectures, ResNet-18 was selected as the optimal model. The avgpool layer of the ResNet-18 model was extracted as raw deep learning features and subjected to PCA for dimensionality reduction, yielding a final set of 20 deep learning features. For the deep-radiomics model, we first fused bioinformatics features with deep learning features. After feature dimensionality reduction of the fused features, we ultimately obtained 10 joint deep-bioinformatics features (Figure 2). We then used the SHAP method to perform an interpretability analysis on the bioinformatics features ultimately included in the model (Figure 3).
Model development
In this study, we developed a radiomics model and a deep-radiomics model. Based on comparisons of the performance of different machine learning classifiers across the models, we selected MLP for the radiomics model and ExtraTrees for the deep-radiomics model. Univariate LR analysis showed that none of the clinical variables were significantly associated with the outcome. When combined with baseline data comparisons, none of the clinical indicators showed differences between groups; therefore, it is not possible to establish a reliable clinical predictive model based on the available clinical parameters. The AUC values for both models on the training, test, and external validation sets are shown in Table 3.
Table 3
| Cohort | Model type | AUC (95% CI) | Sensitivity | Specificity | Accuracy | PPV | NPV |
|---|---|---|---|---|---|---|---|
| Training cohort | Radiomics model | 0.777 (0.701–0.854) | 0.778 | 0.684 | 0.710 | 0.486 | 0.889 |
| Deep-radiomics model | 0.881 (0.821–0.940) | 0.911 | 0.736 | 0.784 | 0.569 | 0.911 | |
| Test cohort | Radiomics model | 0.703 (0.528–0.877) | 0.917 | 0.500 | 0.619 | 0.423 | 0.917 |
| Deep-radiomics model | 0.853 (0.708–0.997) | 0.667 | 0.967 | 0.881 | 0.889 | 0.667 | |
| External validation cohort | Radiomics model | 0.502 (0.350–0.695) | 0.769 | 0.429 | 0.493 | 0.238 | 0.769 |
| Deep-radiomics model | 0.802 (0.673–0.931) | 0.846 | 0.750 | 0.768 | 0.440 | 0.846 |
AUC, area under the receiver operating characteristic curve; CI, confidence interval; NPV, negative predictive value; PPV, positive predictive value.
Model effectiveness and explainability
The deep-radiomics model significantly outperformed the radiomics model; the AUC values for the two models are shown in Figure 4. ROC curve analysis indicated that the deep-radiomics model yielded a higher clinical net benefit than the radiomics model (Figure 5). In the external validation set, the Delong test revealed that the difference in predictive performance between the two models was highly statistically significant (P<0.02), suggesting that the deep-radiomics model has greater clinical utility. Results from the external validation of the calibration curve and Brier score (Figure 6) show that the Brier score for the combined deep learning–bioinformatics model was 0.125, which is lower than that of the bioinformatics model (0.153). This suggests that this model exhibits smaller overall error between predicted probabilities and actual disease risk and demonstrates superior calibration performance. Additionally, SHAP model prediction waterfall plots and force maps were introduced to provide individual-level explanations of the basis and reasoning process for predicting the risk of distant metastasis in a representative case from the external validation set (Figure 7).
Discussion
This study developed predictive models based on radiomic and deep learning features of rectal tumors and the rectal mesentery to assess their value in predicting EP within 2 years post-treatment in patients with rectal cancer. The results showed that, among various machine learning classifiers, the radiomic model ultimately selected a MLP as the classifier for model construction, while the deep-radiomic hybrid model selected ExtraTrees as the optimal classifier. In the validation set, the radiomics model’s predictive performance was relatively limited (AUC =0.502), whereas the deep-radiomics model demonstrated better predictive performance (AUC =0.802). The predictive performance of the radiomics model developed in this study needs further improvement, as its performance on the external validation set was significantly lower than that on the training and test sets. The primary reason is that traditional radiomics features (shape, grayscale, and texture) are highly sensitive to variations in scanner models, slice thickness, reconstruction algorithms, and imaging parameters, resulting in inherently weak cross-center robustness and reproducibility. First, there is heterogeneity in device parameters and scanning protocols between the external validation cohort and the training set, which can easily lead to systematic shifts in the grayscale distribution and numerical baselines of the omics features, thereby disrupting the original associations between the features and prognostic outcomes. Second, the relatively limited sample size of the external validation set further amplifies the interference caused by population heterogeneity on the model’s predictive performance. These factors collectively result in a significant decline in the model’s generalization ability in the independent external cohort, ultimately leading to validation performance falling below the threshold for valid prediction (AUC of 0.502).
Compared with radiomics-only models, the deep-radiomic hybrid model developed in this study displayed superior performance in predicting the EP of rectal cancer. This result may be attributed to deep learning’s ability to capture more complex imaging information. Studies have shown that incorporating deep learning features may provide richer imaging information (21,22), thereby improving the model’s ability to predict the risk of EP of rectal cancer. Conventional radiomics features are largely based on predefined morphological and textural descriptors; although they can reflect the imaging phenotype of tumors to a certain extent, they remain limited in capturing complex spatial structures and higher-order features (23,24). In contrast, deep learning models can automatically learn latent representations from imaging data using multi-layer neural networks, thereby extracting higher-dimensional feature information. When deep learning features are fused with traditional radiomic features, they can complement each other, thereby enhancing the model’s predictive effectiveness. Furthermore, this study includes imaging information from both the rectal tumor and the rectal mesentery for analysis. Previous studies have shown that the rectal mesentery is not only a crucial anatomical structure for local tumor invasion and lymphatic metastasis, but that its imaging characteristics may also reflect the tumor microenvironment and invasive behavior (25,26). Therefore, a complete evaluation of the imaging phenotypes of both the tumor and the surrounding tissues may result in a more accurate and holistic assessment of disease progression risk in patients with rectal cancer.
Previous studies have explored the use of radiomics or deep learning approaches to predict the risk of recurrence in rectal cancer (27-29). Among these, several multicenter studies have reported that fusion models obtain relatively high accuracy in predicting early recurrence, with AUC values of approximately 0.863–0.880 (11). The study by Jin et al. (30) also used only T2-weighted cross-sectional images and included both the tumor and surrounding tissues to predict prognosis in patients with rectal cancer. However, most of these studies used recurrence or disease-free survival as endpoints, with relatively little focus on the risk of EP after treatment. The study by Li et al. (11) compared the predictive value of four types of models—clinical models, traditional radiomics models, deep learning models, and radiomics-deep learning fusion models—for early recurrence of rectal cancer; however, it focused only on the tumor itself rather than the surrounding tissues and did not provide interpretable insights at the individual level. The present study used EP as the predictive endpoint, further integrating deep learning features with radiomic features, and incorporating information from the rectal mesenteric region into the analysis, achieving good predictive performance (AUC =0.802) in the validation cohort. Compared with previous studies, this study expands both the predictive endpoint and the scope of imaging analysis, suggesting that an extensive examination of imaging features from both the tumor and its surrounding tissues may facilitate earlier identification of high-risk patients.
To further explicate the model’s predictive mechanism, this study used the SHAP method to evaluate the contribution of each feature to the model’s predictive performance. The SHAP analysis showed that deep learning features played an important role in the model’s predictions, with DL_resnet18_3 and DL_resnet18_6 ranking as the two most influential features. These features were extracted from the ResNet-18 network and can capture complex imaging patterns that are difficult to characterize with conventional radiomics features. Furthermore, several radiomic features derived from the rectal mesenteric region (T2 surround) also exhibit high importance in the model, including wavelet-based texture and first-order statistical features (31,32). These features may reflect changes in the tumor microenvironment (33). The mesorectal region, an important anatomical site for local tumor invasion and lymphatic metastasis, may exhibit imaging changes closely associated with tumor aggressiveness and disease progression. Consequently, this study’s data show that combining deep learning features with imaging information from surrounding tissue can provide a more comprehensive representation of the tumor’s biological behavior, thereby enhancing the model’s predictive ability. It is worth noting that SHAP analysis not only identifies key features influencing model predictions but also explains the model’s decision-making process at the individual level. Taking a typical case as an example (Figure 7A), the baseline predicted risk was 0.27; under the combined influence of multiple high-risk features, the final predicted risk rose to 0.339, indicating that the patient’s risk of EP was higher than the overall average. These results demonstrate that this model not only outputs risk predictions but also identifies the key radiological features driving the increased risk, thereby supporting clinicians in understanding the basis for the model’s predictions.
From a clinical translation perspective, the predictive probabilities generated by the model in this study can be used for preoperative risk stratification. Based on the threshold of the deep learning-based radiomics model, patients with a predicted probability ≥0.265 can be classified as being at high risk for early disease progression. For such patients, clinicians may consider adopting more aggressive management strategies, including appropriately shortening postoperative follow-up intervals and intensifying imaging monitoring of the thoracic, abdominal, and pelvic regions to improve the early detection rate of recurrent and metastatic lesions. For patients with a lower predicted risk, this approach may help avoid excessive testing and the waste of medical resources. It should be noted that this study has not yet established a prospectively validated clinical intervention threshold; therefore, the model is currently more suitable as an auxiliary decision-making tool, and its clinical value still requires further validation in large-scale prospective studies.
This study still has several limitations. Firstly, as a retrospective study, it is subject to potential selection bias, and the sample size was relatively limited. Although a validation cohort was included, the model’s generalizability requires more validation in larger, multicenter prospective studies. Secondly, the model was primarily constructed using radiomic and deep learning features, without including clinical factors or molecular biological indicators, which may limit further improvements in the model’s performance. Thirdly, the image segmentation process relies on manual operations; despite the adoption of a standardized workflow, this may still introduce inter-observer variability. Finally, although the SHAP method was used to interpret the model in this study, the underlying biological significance of the deep learning features remains to be further investigated to strengthen the model’s interpretability and clinical usefulness. Looking forward, several avenues for future research are warranted. Prospective, multicenter studies should be conducted to further validate and generalize the predictive model in diverse patient populations and clinical settings. Integrating relevant clinical variables and molecular or genomic data into future models could enhance predictive accuracy and facilitate more personalized risk assessment. The development and implementation of automated or semi-automated segmentation methods may also help reduce observer variability and improve workflow efficiency. Additionally, exploring approaches to clarify the biological meaning of deep learning features, such as correlating them with histopathologic or molecular findings, could further strengthen the interpretability and clinical applicability of these predictive models. At the same time, intelligent image analysis technology must not only enable retrospective prognostic modeling but also evolve toward real-time intraoperative navigation and integration into the surgical workflow, thereby enhancing the precision of intraoperative diagnosis and treatment in minimally invasive surgery (34).
Conclusions
This study integrated imaging features of the primary tumor and the rectal mesentery to develop a deep radiomics model. Following internal cross-validation and external cohort validation, the model demonstrated good predictive performance and can support preoperative risk stratification for high-risk rectal cancer patients.
Acknowledgments
None.
Footnote
Reporting Checklist: The authors have completed the TRIPOD reporting checklist. Available at https://jgo.amegroups.com/article/view/10.21037/jgo-2026-0451/rc
Data Sharing Statement: Available at https://jgo.amegroups.com/article/view/10.21037/jgo-2026-0451/dss
Peer Review File: Available at https://jgo.amegroups.com/article/view/10.21037/jgo-2026-0451/prf
Funding: None.
Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://jgo.amegroups.com/article/view/10.21037/jgo-2026-0451/coif). The authors have no conflicts of interest to declare.
Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. The study was approved by the Ethics Committees of The First Affiliated Hospital of Soochow University [approval No. (2025) Ethical Review No. 739] and the Suzhou Hospital Affiliated to Nanjing Medical University (approval No. K-2025-206-K01). Because this was a retrospective study, patients did not need to sign informed consent forms.
Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.
References
- Siegel RL, Miller KD, Wagle NS, et al. Cancer statistics, 2023. CA Cancer J Clin 2023;73:17-48. [Crossref] [PubMed]
- Matsuda T, Fujimoto A, Igarashi Y. Colorectal Cancer: Epidemiology, Risk Factors, and Public Health Strategies. Digestion 2025;106:91-9. [Crossref] [PubMed]
- Zhang Y, Yang Z, Feng Y, et al. Development and Validation of Nomogram Models Incorporating the Inflammatory Nutritional Index CALLY for Predicting Survival in Locally Advanced Rectal Cancer After Neoadjuvant Chemoradiotherapy. Cancer Manag Res 2025;17:2961-75. [Crossref] [PubMed]
- Wang J, Gao QK, Wang YF, et al. A new insights of mesorectum. Zhonghua Wei Chang Wai Ke Za Zhi 2022;25:321-6. [Crossref] [PubMed]
- Dumont F, Barreteau T, Simon G, et al. Laparoscopic Extraperitoneal Approach to Bilateral Pelvic Lymph Node Dissection in Low Rectal Cancer: Technique with Video and 3D Modeling. Ann Surg Oncol 2022;29:109-11. [Crossref] [PubMed]
- Handa P, Goel N, Indu S, et al. AI-KODA score application for cleanliness assessment in video capsule endoscopy frames. Minim Invasive Ther Allied Technol 2024;33:311-20. [Crossref] [PubMed]
- Qiu C, Xia Y, Feng Z, et al. The Diagnostic Value of Deep Learning for Multi-Classification of Rectal Cancer T Staging Based on Regional Attention. Diagnostics (Basel) 2026;16:1525. [Crossref] [PubMed]
- Yao X, Deng S, Han X, et al. Deep Learning Algorithm Based MRI Radiomics and Pathomics for Predicting Microsatellite Instability Status in Rectal Cancer: A Multicenter Study. Acad Radiol 2025;32:1934-45. [Crossref] [PubMed]
- Lu H, Yuan Y, Liu M, et al. Predicting pathological complete response following neoadjuvant chemoradiotherapy (nCRT) in patients with locally advanced rectal cancer using merged model integrating MRI-based radiomics and deep learning data. BMC Med Imaging 2024;24:289. [Crossref] [PubMed]
- Yang Y, Han K, Xu Z, et al. Development and Validation of Multiparametric MRI-based Interpretable Deep Learning Radiomics Fusion Model for Predicting Lymph Node Metastasis and Prognosis in Rectal Cancer: A Two-center Study. Acad Radiol 2025;32:2642-54. [Crossref] [PubMed]
- Li Z, Qin Y, Liao X, et al. Comparison of clinical, radiomics, deep learning, and fusion models for predicting early recurrence in locally advanced rectal cancer based on multiparametric MRI: a multicenter study. Eur J Radiol 2025;189:112173. [Crossref] [PubMed]
- He J, Tan X, Lin H, et al. Magnetic resonance imaging-based radiomics of mesorectum for predicting extramural venous invasion in patients with rectal cancer: a bi-centric study. Cancer Imaging 2026;26:69. [Crossref] [PubMed]
- Ye Y, Zhao K, Feng L, et al. Higher local recurrence as the distinct failure pattern in early-onset rectal cancer: a tailored MRI score to guide therapy. NPJ Precis Oncol 2025;10:44. [Crossref] [PubMed]
- Liang ZY, Yu ML, Yang H, et al. Beyond the tumor region: Peritumoral radiomics enhances prognostic accuracy in locally advanced rectal cancer. World J Gastroenterol 2025;31:99036. [Crossref] [PubMed]
- Qin S, Lu S, Liu K, et al. Radiomics from Mesorectal Blood Vessels and Lymph Nodes: A Novel Prognostic Predictor for Rectal Cancer with Neoadjuvant Therapy. Diagnostics (Basel) 2023;13:1987. [Crossref] [PubMed]
- Deng B, Wang Q, Liu Y, et al. A nomogram based on MRI radiomics features of mesorectal fat for diagnosing T2- and T3-stage rectal cancer. Abdom Radiol (NY) 2024;49:1850-60. [Crossref] [PubMed]
- Fernandes MC, Gollub MJ, Brown G. The importance of MRI for rectal cancer evaluation. Surg Oncol 2022;43:101739. [Crossref] [PubMed]
- Zhan Z, Chen B, Cheng H, et al. Identification of prognostic signatures in remnant gastric cancer through an interpretable risk model based on machine learning: a multicenter cohort study. BMC Cancer 2024;24:547. [Crossref] [PubMed]
- Igami T, Hayashi Y, Yokyama Y, et al. Development of real-time navigation system for laparoscopic hepatectomy using magnetic micro sensor. Minim Invasive Ther Allied Technol 2024;33:129-39. [Crossref] [PubMed]
- Safari M, Mahmoudi L, Baker EK, et al. Recurrence and Postoperative Death in Patients with Colorectal Cancer: A New Perspective via Semi-competing Risk Framework. Turk J Gastroenterol 2023;34:736-46. [Crossref] [PubMed]
- Can Z, Aydin E. Explainable CNN-Radiomics Fusion and Ensemble Learning for Multimodal Lesion Classification in Dental Radiographs. Diagnostics (Basel) 2025;15:1997. [Crossref] [PubMed]
- Li J, Zhou Y, Wang X, et al. An MRI-based multi-objective radiomics model predicts lymph node status in patients with rectal cancer. Abdom Radiol (NY) 2021;46:1816-24. [Crossref] [PubMed]
- Tran AT, Wen J, Abou Karam G, et al. Comparing Handcrafted Radiomics Versus Latent Deep Learning Features of Admission Head CT for Hemorrhagic Stroke Outcome Prediction. BioTech (Basel) 2025;14:87. [Crossref] [PubMed]
- Liu Y, Pan Y, Wang Q, et al. Integration of intratumoral/peritumoral radiomics and deep learning for predicting overall survival in non-small cell lung cancer patients: a multicenter study. Front Oncol 2025;15:1669200. [Crossref] [PubMed]
- Li H, Chen XL, Liu H, et al. MRI-based multiregional radiomics for predicting lymph nodes status and prognosis in patients with resectable rectal cancer. Front Oncol 2022;12:1087882. [Crossref] [PubMed]
- Chiloiro G, Cusumano D, Romano A, et al. Delta Radiomic Analysis of Mesorectum to Predict Treatment Response and Prognosis in Locally Advanced Rectal Cancer. Cancers (Basel) 2023;15:3082. [Crossref] [PubMed]
- Xie PY, Zeng ZM, Li ZH, et al. MRI-based radiomics for stratifying recurrence risk of early-onset rectal cancer: a multicenter study. ESMO Open 2024;9:103735. [Crossref] [PubMed]
- Fu S, Xia T, Li Z, et al. Baseline MRI-based radiomics improving the recurrence risk stratification in rectal cancer patients with negative carcinoembryonic antigen: A multicenter cohort study. Eur J Radiol 2025;182:111839. [Crossref] [PubMed]
- Jiang X, Zhao H, Saldanha OL, et al. An MRI Deep Learning Model Predicts Outcome in Rectal Cancer. Radiology 2023;307:e222223. [Crossref] [PubMed]
- Jin X, Xu J, Hu Y, et al. Macro Habitat-Based T2-Weighted MRI Radiomics and Deep Learning Fusion for Predicting Treatment Response and Prognosis After Neoadjuvant Chemoradiotherapy in Locally Advanced Rectal Cancer. Cancer Med 2026;15:e71773. Erratum in: Cancer Med 2026;15:e72056.
- Fan S, Wang N, Wen Y, et al. Predictive model based on mesorectal fat radiomics for pathological complete response to neoadjuvant chemoradiotherapy in locally advanced rectal cancer. Eur J Radiol 2025;192:112378. Erratum in: Eur J Radiol 2025;193:112421.
- Jayaprakasam VS, Paroder V, Gibbs P, et al. MRI radiomics features of mesorectal fat can predict response to neoadjuvant chemoradiation therapy and tumor recurrence in patients with locally advanced rectal cancer. Eur Radiol 2022;32:971-80. [Crossref] [PubMed]
- Treillard S, Schwob R, Mouysset S, et al. Biological feature-based machine learning in histopathological images: a systematic review. J Pathol Inform 2026;20:100539. [Crossref] [PubMed]
- Boland PA, McEntee PD, Cucek J, et al. Protocol for CLASSICA software as medical device trial. Minim Invasive Ther Allied Technol 2025;34:441-6. [Crossref] [PubMed]


