EndoFAst3r: A Self-Supervised Framework for Monocular Depth and Pose Estimation in Endoscopic Surgery

Tuesday 08 April 2025


A team of researchers has made a significant breakthrough in developing a new system for estimating depth and camera pose in robotic-assisted surgery. This technology, known as Endo-FASt3r, uses self-supervised learning to adapt foundation models to the specific challenges of endoscopic imaging.


In traditional computer vision, depth estimation is typically done using stereo cameras or structured light. However, these methods are not suitable for endoscopy due to the complex and dynamic nature of the surgical environment. Endo-FASt3r addresses this challenge by leveraging a self-supervised learning approach that can adapt to the unique conditions of robotic-assisted surgery.


The system consists of two main components: a depth module that estimates the distance between objects in the scene, and a pose module that calculates the camera’s movement relative to the patient. Both modules are trained using a novel adaptation technique called DoMoRA, which combines the benefits of low-rank and full-rank updates for faster convergence.


The researchers tested Endo-FASt3r on three publicly available datasets: SCARED, Hamlyn, and StereoMIS. The results show that the system outperforms existing methods in both depth estimation and camera pose estimation tasks. In particular, Endo-FASt3r achieved a substantial improvement of 7% to 10% in the average translation error (ATE) metric compared to the second-best approach.


The significance of this achievement lies not only in its technical prowess but also in its potential impact on patient care. Robotic-assisted surgery offers numerous benefits, including reduced recovery time and improved accuracy. However, it requires precise control over the surgical instruments, which can be challenging due to the limited visibility and complex anatomy of the surgical site.


Endo-FASt3r’s ability to accurately estimate depth and camera pose can significantly enhance the performance of robotic-assisted surgery systems. By providing a more accurate understanding of the surgical environment, surgeons can make more informed decisions and perform procedures with greater precision and confidence.


The researchers plan to further refine Endo-FASt3r by exploring its application to other areas of computer vision, such as 3D reconstruction and tracking in medical imaging. With continued advancements in this field, we may see the development of more sophisticated surgical robots that can improve patient outcomes and revolutionize the way we approach minimally invasive procedures.


Cite this article: “EndoFAst3r: A Self-Supervised Framework for Monocular Depth and Pose Estimation in Endoscopic Surgery”, The Science Archive, 2025.


Robotics, Computer Vision, Endoscopy, Depth Estimation, Camera Pose, Self-Supervised Learning, Robotic-Assisted Surgery, Medical Imaging, 3D Reconstruction, Tracking.


Reference: Mona Sheikh Zeinoddin, Mobarakol Islam, Zafer Tandogdu, Greg Shaw, Mathew J. Clarkson, Evangelos Mazomenos, Danail Stoyanov, “Endo-FASt3r: Endoscopic Foundation model Adaptation for Structure from motion” (2025).


Leave a Reply