Surgical Video Understanding
Online, causal recognition of surgical workflow — from single frames to multi-hour procedures. Long-video transformers, key-information compression and stable temporal inference.
Medical AI · Surgical Video Understanding · Vision-Language Models
I received my PhD from King's College London (School of Biomedical Engineering & Imaging Sciences) in January 2026, supervised by Prof. Sébastien Ourselin (primary supervisor), Prof. Prokar Dasgupta and Dr. Alejandro Granados. I received my Master's degree from Huazhong University of Science and Technology in 2021, advised by Prof. Xiang Bai.
My research builds AI systems that understand surgical video in real time — online phase recognition over multi-hour procedures, streaming vision-language models for the operating room, and label-efficient segmentation of instruments and anatomy. I also work on broader medical imaging problems (echocardiography, stroke lesion segmentation, structural MRI) and on vision-language learning with weak supervision.
From January to July 2025 I worked part-time as a Data Scientist at Proximie, bringing surgical-video models closer to clinical deployment.
Online, causal recognition of surgical workflow — from single frames to multi-hour procedures. Long-video transformers, key-information compression and stable temporal inference.
Streaming VLMs for the operating room, weakly supervised referring comprehension & segmentation, and noise-injected cross-modal alignment for language-driven visual tasks.
Unsupervised surgical instrument segmentation from low-quality optical flow, fast boundary-to-pixel direction segmentation, and diffusion-based crack segmentation.
Training-free cardiac phase detection in echocardiography, ischemic stroke lesion segmentation, and gray-matter-guided attention for Alzheimer's diagnosis from structural MRI.
* equal contribution · † corresponding author
Surgical video understanding, streaming vision-language models, label-efficient medical segmentation. Primary supervisor: Prof. Sébastien Ourselin; co-supervised by Prof. Prokar Dasgupta and Dr. Alejandro Granados.
Surgical-video machine learning for a clinical telepresence and analytics platform.
Applied computer vision for education products.
Image segmentation and image restoration; publications at CVPR 2020 and WACV 2021.
Open to collaboration on surgical AI, medical video understanding and vision-language models.