Rectal cancer accounts for roughly one-third to one-half of all colorectal cancers in China, which sees approximately 496,000 new colorectal cancer cases every year. The standard surgical treatment for rectal cancer is a procedure called total mesorectal excision (TME), first described by surgeon Richard Heald in 1982.
TME involves removing the rectum along with the surrounding fatty tissue envelope called the mesorectum, which contains lymph nodes where cancer can spread. When performed correctly, TME dramatically reduces the chance that cancer will come back - contributing to a 5-year survival rate of approximately 70% for rectal cancer patients treated with surgery plus chemotherapy and radiation.
The key challenge in TME is identifying the correct dissection plane - the precise tissue layer between the mesorectum and the surrounding pelvic structures. If the surgeon cuts too deep or too shallow, they risk incomplete cancer removal, damage to the pelvic autonomic nerves (which control urinary and sexual function), injury to major blood vessels, or accidental damage to the ureters (the tubes connecting the kidneys to the bladder).
Ureteral injuries occur in 0.7% to 6% of rectal cancer operations and can have serious long-term consequences for the patient's quality of life. Even with modern tools like 3D laparoscopy and robotic surgical systems, factors such as obesity, bleeding, tumor adhesion, and narrow pelvic anatomy can obscure the surgical field and increase the risk of these complications.
Artificial intelligence (AI) has shown remarkable ability to detect patterns in images and videos that are difficult for the human eye to reliably identify in real time. In medicine, AI has already been applied to radiology, pathology, and endoscopy - but its application during live surgery represents a newer and potentially transformative frontier.
In colorectal surgery specifically, researchers have begun using AI to recognize surgical phases, identify fascial planes, and locate critical anatomical structures. However, prior systems were limited - they recognized only one or two structures at a time, and their accuracy was insufficient for routine clinical use.
A comprehensive AI navigation model - one that simultaneously recognizes multiple critical structures including arteries, veins, ureters, fascial layers, and pelvic nerves - could act like a real-time GPS overlay for the surgical field. It would alert surgeons to the location of at-risk structures before they are accidentally damaged.
This study set out to build exactly such a system, using a cutting-edge AI model called SAM2 (Segment Anything in Images and Video), developed by Meta AI. SAM2 was the first model of its kind applied to anatomical structure recognition in rectal cancer surgery, offering particular strengths in real-time video segmentation - the ability to continuously track and label structures as the surgical video streams live.
The research team began by building a comprehensive database of intraoperative images from rectal cancer surgeries performed at their institution between January 2016 and April 2024. They collected high-definition screenshots and video frames from laparoscopic and robotic (Da Vinci) procedures performed on patients with confirmed rectal cancer.
A total of 6,700 high-quality images from 325 patients were selected after careful quality screening. Images were checked for sharpness using a mathematical measure called the Laplacian variance - blurry images were excluded. All images were standardized to a resolution of 1280 by 720 pixels. The images were categorized into four groups: pelvic autonomic nerve views (2,050 images), arteries and veins (2,230 images), fascial layers and excision planes (2,800 images), and ureter images (1,800 images).
Expert annotation was performed by two gastrointestinal surgeons with more than 5 years of experience, with disagreements resolved by a professor with over 20 years of experience. This precise manual annotation created the ground-truth labels the AI model would learn from. Inter- and intra-observer agreement was formally assessed using the kappa concordance test to ensure reliability.
The SAM2 model architecture includes seven key components: an Image Encoder, Prompt Encoder, Memory Attention module, Mask Decoder, Memory Encoder, Memory Bank, and Geometric Feature extractor. Training ran for 1,000 iterations, with data augmentation techniques such as rotation, mirroring, and blurring applied to prevent the model from simply memorizing the training examples.
The model was evaluated using several standard image segmentation performance metrics. The most important were mean Intersection over Union (mIoU), which measures how well the model's predicted regions overlap with the expert-annotated ground truth; precision, which measures how often the model's identifications are correct; recall, which measures how many actual structures the model successfully detected; and F1 score, which combines precision and recall into a single balanced measure.
To test whether the model could perform reliably in different hospitals and patient populations, the team conducted two external validation studies. The first used 687 images from 43 patients at a separate institution; the second used 155 images from 19 patients at yet another site. Performance in these external datasets was compared against the results from the original training institution.
The model was also directly compared against junior surgeons (those with less than 10 years of clinical experience) to quantify how AI performance compares to human expertise at that career stage. A senior professor evaluated the annotations from both the AI model and the junior surgeons in a blinded manner - meaning the professor did not know which annotations came from which source.
Finally, the integrated model was tested on live surgical videos and deployed in an actual operating room setting - integrated into a dual-screen 3D laparoscopic display and as a picture-in-picture overlay within the Da Vinci Xi robotic surgery system. This real-world clinical deployment tested the system under the full complexity of live surgical conditions.
The trained model achieved strong performance across most anatomical structures. The highest accuracy was seen for the Toldt fascia (the key tissue plane that must be preserved during TME), with an mIoU of 0.8985 and F1 score of 0.9466. Arteries (mIoU 0.8095, F1 0.8947) and veins (mIoU 0.8114, F1 0.8959) were also identified with high accuracy.
The ureter was identified with an mIoU of 0.7086 and F1 score of 0.8295 - solid performance for a structure that is notoriously difficult to locate, particularly in cases with adhesion or bleeding. The separation layer (the correct dissection plane) achieved an mIoU of 0.7281 and F1 of 0.8427. The pelvic autonomic nerves were the most challenging structure, with lower scores reflecting their small size and variable anatomy.
In external validation at two separate hospitals, the model showed broadly consistent performance for the main structures. However, nerve and reproductive blood vessel recognition dropped significantly - likely reflecting differences in image characteristics across surgical teams and equipment. When just 10% of external validation images were added to the training set, nerve recognition improved markedly, demonstrating that the model can be quickly customized for new hospitals with minimal additional data.
Compared directly to junior surgeons, the AI model outperformed human physicians on ureter, reproductive vessel, and autonomic nerve identification. The model recognized all structures in an external validation dataset in just 63 seconds, whereas the same task took an expert professor 36 minutes - a roughly 34-fold speed advantage that is clinically meaningful in a real-time surgical setting.
The truly novel aspect of this study was not just building an accurate model in a laboratory setting, but deploying it during actual surgery. The AI recognition system was integrated into both a 3D laparoscopic display (using a KARL STORZ system) and the Da Vinci Xi robotic surgery system as a picture-in-picture overlay at a major Chinese hospital.
In the operating room, each anatomical structure was displayed in a distinct color: arteries in red, veins in blue, ureters and fascial planes in other designated colors. This color-coded overlay gave the surgical team an instant visual reference for where each critical structure was located, without requiring them to interpret raw AI data or perform additional mental calculations.
One particularly impressive demonstration occurred in a case where the inferior mesenteric artery was not yet fully exposed. Despite the incomplete visibility, the model correctly identified all three of its branches - the left colic artery, sigmoid colon artery, and superior rectal artery - and highlighted them before they were fully visible to the operating surgeon. This predictive navigation capability goes beyond simple recognition to anticipate what lies just outside the current field of view.
Even in cases with limited surgical visibility due to bleeding, tissue smoke, or traction, the ureter was recognized accurately in real time. This is precisely the scenario where ureteral injuries occur - when the field is obscured and the surgeon is working under time pressure. The AI provided a critical safety reminder that prevented potential damage.
Earlier AI systems for rectal cancer surgery typically focused on a single structure or a narrow subset of the anatomy. The SegFormer-based model from Kolbinger et al. - one of the most prominent prior systems - achieved F1 scores of 0.60 for arteries, 0.65 for veins, and 0.58 for ureters. The current SAM2-based model substantially improved on these figures, achieving 0.8947 for arteries, 0.8959 for veins, and 0.8295 for ureters.
The current study also improved substantially over Kolbinger's fascial plane recognition, which achieved an F1 of 0.78 for the Toldt fascia. The SAM2 model reached 0.9466 - a meaningful gain for the single most important landmark in determining correct dissection plane during TME.
Crucially, this model is the first to simultaneously recognize all major TME structures in a single integrated system. Previous models were siloed - recognizing either vessels, or fascial planes, or ureters, but not all together. Integration matters because in the operating room, surgeons need to understand the spatial relationships between structures, not just identify each in isolation.
The use of the SAM2 architecture was particularly advantageous because of its memory-based video segmentation capabilities. SAM2 can track structures across video frames even when they temporarily disappear from view, and its training on an enormous general dataset (SA-1B, with over 1 billion masks) gives it strong generalization ability. The mask autoencoder (MAE) used in training specifically helps the model learn spatial relationships between nearby structures - essential for predicting what lies just beyond the current surgical field.
This study represents a significant step toward making rectal cancer surgery safer and more standardized, particularly in settings where experienced surgeons are not always available. For surgical teams with limited experience, an AI navigation overlay can serve as a real-time guide - helping less experienced surgeons identify critical planes, avoid injury-prone structures, and maintain oncologic quality.
The model's demonstrated ability to be quickly customized for new hospitals with just 10% additional local training data is an important practical feature. Healthcare systems rarely have identical equipment, lighting, patient populations, or surgical styles, so models that adapt efficiently are far more likely to achieve widespread adoption than those requiring full retraining from scratch.
The authors acknowledge several limitations. The study was conducted primarily at a single center, with external validation providing only partial evidence of broader generalizability. Annotation bias remains possible despite quality controls. Most importantly, the study has not yet demonstrated that using the AI navigation system actually reduces complication rates or improves patient outcomes in a controlled clinical trial - that evidence will require prospective studies.
Looking ahead, the authors call for prospective multi-center studies that directly measure surgical complications - particularly ureteral injuries and incomplete mesorectal excision rates - in patients whose surgeries used the AI system versus those who did not. If such trials confirm the benefits suggested by this technical validation work, AI-guided TME surgery could become a new standard of care for rectal cancer treatment worldwide.