(2) Hassan Silkan
(3) Insaf Bellamine
*corresponding author
AbstractFine-grained surgical action recognition in laparoscopic videos remains a challenge, even with recent advances in deep learning. While current VideoMAE approaches reach 89.11% accuracy on cholecystectomy tasks, they face specific limitations. Random masking strategies often miss surgical instruments that occupy only 10% to 15% of frames. Furthermore, context-independent models struggle with visually similar actions across different phases, and symmetric two-stream architectures tend to waste computational resources. To solve this, we developed SA-VideoMAE, a surgical-aware video masked autoencoder specifically designed for laparoscopic action recognition. Our method utilizes surgical-aware adaptive masking that integrates YOLOv7x object detection to prioritize instrument patches. This increased instrument visibility from 10% to 60% during training, ensuring the model focuses on action-relevant regions rather than static backgrounds. We also utilized phase-conditioned hierarchical attention to inject learnable phase embeddings into the attention mechanisms, enabling the model to disambiguate visually similar actions based on surgical context. For efficiency, our asymmetric dual-stream architecture processes RGB using ViT-Base (86M parameters) and optical flow using ViT-Tiny (5.7M parameters), achieving a 47% parameter reduction compared to symmetric designs. Our training process then balanced reconstruction, classification, temporal consistency, and phase prediction through a novel multi-objective optimization strategy. Results from Cholec80's Calot's Triangle Dissection phase show 93.5% accuracy, representing a 4.4 percentage-point improvement over the verified baseline. Notably, challenging action recall improved from 51% to 74% while maintaining real-time inference at 62ms per clip. These findings demonstrate that encoding surgical domain knowledge into video architectures significantly enhances action recognition performance.
Keywordssurgical action recognition; masked autoencoders; phase-conditioned attention; asymmetric dual-stream; laparoscopic cholecystectomy
|
DOIhttps://doi.org/10.26555/ijain.v12i2.2357 |
Article metricsAbstract views : 123 | PDF views : 2 |
Cite |
Full Text Download
|
References
[1] K. He, X. Chen, S. Xie, Y. Li, P. Dollar, and R. Girshick, “Masked Autoencoders Are Scalable Vision Learners,” in 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Jun. 2022, pp. 15979–15988. doi: 10.1109/CVPR52688.2022.01553.
[2] Z. Tong, Y. Song, J. Wang, and L. Wang, “VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training,” in Advances in Neural Information Processing Systems, Neural information processing systems foundation, Oct. 2022. Accessed: Feb. 16, 2026. [Online]. Available: http://arxiv.org/abs/2203.12602.
[3] H.-H. Yen et al., “Automated surgical action recognition and competency assessment in laparoscopic cholecystectomy: a proof-of-concept study,” Surg. Endosc., vol. 39, no. 5, pp. 3006–3016, May 2025, doi: 10.1007/s00464-025-11663-y.
[4] S. Yang, L. Luo, Q. Wang, and H. Chen, “Surgformer: Surgical Transformer with Hierarchical Temporal Attention for Surgical Phase Recognition,” in Medical Image Computing and Computer Assisted Intervention – MICCAI 2024, Springer, Cham, 2024, pp. 606–616. doi: 10.1007/978-3-031-72089-5_57.
[5] Y. Liu et al., “SKiT: a Fast Key Information Video Transformer for Online Surgical Phase Recognition,” in 2023 IEEE/CVF International Conference on Computer Vision (ICCV), IEEE, Oct. 2023, pp. 21017–21027. doi: 10.1109/ICCV51070.2023.01927.
[6] M. C. Schiappa, Y. S. Rawat, and M. Shah, “Self-Supervised Learning for Videos: A Survey,” ACM Comput. Surv., vol. 55, no. 13s, pp. 1–37, Dec. 2023, doi: 10.1145/3577925.
[7] W. G. C. Bandara, N. Patel, A. Gholami, M. Nikkhah, M. Agrawal, and V. M. Patel, “AdaMAE: Adaptive Masking for Efficient Spatiotemporal Learning with Masked Autoencoders,” in 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Jun. 2023, pp. 14507–14517. doi: 10.1109/CVPR52729.2023.01394.
[8] M. L. Mostafa et al., “Surgical Flow Masked Autoencoder for Event Recognition,” in MIDL 2025 Conference Proceedings of Machine Learning Research, M. L. Mostafa, ~Mayar_Lotfy_Mostafa1, and N. N. , Anna Alperovich, Dmitrii Fedotov, Ghazal Ghazaei, Stefan Saur, Azade Farshad, Eds., 2025, pp. 1–15. Accessed: Jul. 01, 2026. [Online]. Available: https://openreview.net/forum?id=TAPn4QjbOg¬eId=4fwh1V18nH.
[9] Y. Li et al., “SemiVT-Surge: Semi-supervised Video Transformer for Surgical Phase Recognition,” in Lecture Notes in Computer Science, vol. 15969 LNCS, Springer Science and Business Media Deutschland GmbH, 2026, pp. 478–488. doi: 10.1007/978-3-032-05127-1_46.
[10] R. H. Abiyev, M. Z. Altabel, M. Darwish, and A. Helwan, “A Multimodal Transformer Model for Recognition of Images from Complex Laparoscopic Surgical Videos,” Diagnostics, vol. 14, no. 7, p. 681, Mar. 2024, doi: 10.3390/diagnostics14070681.
[11] F. Shamshad et al., “Transformers in medical imaging: A survey,” Med. Image Anal., vol. 88, no. August, p. 102802, Aug. 2023, doi: 10.1016/j.media.2023.102802.
[12] K. Alomar, H. I. Aysel, and X. Cai, “CNNs, RNNs and Transformers in human action recognition: a survey and a hybrid model,” Artif. Intell. Rev., vol. 58, no. 12, p. 387, Oct. 2025, doi: 10.1007/s10462-025-11388-3.
[13] Z. Liu, K. Chen, S. Wang, Y. Xiao, and G. Zhang, “Deep learning in surgical process modeling: A systematic review of workflow recognition,” J. Biomed. Inform., vol. 162, no. 2, p. 104779, Feb. 2025, doi: 10.1016/j.jbi.2025.104779.
[14] S. N. Gowda, A. Arnab, and J. Huang, “Optimizing Factorized Encoder Models: Time and Memory Reduction for Scalable and Efficient Action Recognition,” in Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), vol. 15068 LNCS, Springer, Cham, 2025, pp. 457–474. doi: 10.1007/978-3-031-72684-2_26.
[15] K. Li et al., “UniFormer: Unifying Convolution and Self-Attention for Visual Recognition,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 45, no. 10, pp. 12581–12600, Oct. 2023, doi: 10.1109/TPAMI.2023.3282631.
[16] Z. Liu et al., “Video Swin Transformer,” in 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Jun. 2022, pp. 3192–3201. doi: 10.1109/CVPR52688.2022.00320.
[17] Y. Liu et al., “LoViT: Long Video Transformer for surgical phase recognition,” Med. Image Anal., vol. 99, no. 6, p. 103366, Jan. 2025, doi: 10.1016/j.media.2024.103366.
[18] B. Ran, B. Huang, S. Liang, and Y. Hou, “Surgical Instrument Detection Algorithm Based on Improved YOLOv7x,” Sensors, vol. 23, no. 11, p. 5037, May 2023, doi: 10.3390/s23115037.
[19] S. Takahashi et al., “Comparison of Vision Transformers and Convolutional Neural Networks in Medical Image Analysis: A Systematic Review,” J. Med. Syst., vol. 48, no. 1, p. 84, Sep. 2024, doi: 10.1007/s10916-024-02105-8.
[20] F. A. Ahmed et al., “Deep learning for surgical instrument recognition and segmentation in robotic-assisted surgeries: a systematic review,” Artif. Intell. Rev., vol. 58, no. 1, p. 1, Nov. 2024, doi: 10.1007/s10462-024-10979-w.
[21] C. I. Nwoye et al., “CholecTriplet2021: A benchmark challenge for surgical action triplet recognition,” Med. Image Anal., vol. 86, no. May, p. 102803, May 2023, doi: 10.1016/j.media.2023.102803.
[22] L. Wang et al., “VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking,” in 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Jun. 2023, pp. 14549–14560. doi: 10.1109/CVPR52729.2023.01398.
[23] M. A. Jamal and O. Mohareri, “SurgMAE: Masked Autoencoders for Long Surgical Video Analysis,” May 2023, p. 14. Accessed: Feb. 16, 2026. [Online]. Available: http://arxiv.org/abs/2305.11451.
[24] N. A. Shah, C. Bandara, S. Sikder, S. S. Vedula, and V. M. Patel, “CSMAE : Cataract Surgical Masked Autoencoder (MAE) Based Pre-Training,” in 2025 IEEE 22nd International Symposium on Biomedical Imaging (ISBI), IEEE, Apr. 2025, pp. 1–5. doi: 10.1109/ISBI60581.2025.10981288.
[25] R. Fujii, M. Hatano, H. Saito, and H. Kajita, “EgoSurgery-Phase: A Dataset of Surgical Phase Recognition from Egocentric Open Surgery Videos,” in Medical Image Computing and Computer Assisted Intervention – MICCAI 2024, Springer, Cham, 2024, pp. 187–196. doi: 10.1007/978-3-031-72089-5_18.
[26] B. Huang, Z. Zhao, G. Zhang, Y. Qiao, and L. Wang, “MGMAE: Motion Guided Masking for Video Masked Autoencoding,” in 2023 IEEE/CVF International Conference on Computer Vision (ICCV), IEEE, Oct. 2023, pp. 13447–13458. doi: 10.1109/ICCV51070.2023.01241.
[27] A. Alfarano, L. Maiano, L. Papa, and I. Amerini, “Estimating optical flow: A comprehensive review of the state of the art,” Comput. Vis. Image Underst., vol. 249, no. December, p. 104160, Dec. 2024, doi: 10.1016/j.cviu.2024.104160.
[28] X. Shi et al., “FlowFormer++: Masked Cost Volume Autoencoding for Pretraining Optical Flow Estimation,” in 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Jun. 2023, pp. 1599–1610. doi: 10.1109/CVPR52729.2023.00160.
[29] Y. Wang, L. Lipson, and J. Deng, “SEA-RAFT: Simple, Efficient, Accurate RAFT for Optical Flow,” in Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), vol. 15065 LNCS, Springer, Cham, 2025, pp. 36–54. doi: 10.1007/978-3-031-72667-5_3.
[30] M. Song, Y. Li, Y. Liu, and L. Yang, “A multitask learning network with interactive fusion for surgical instrument segmentation,” Knowledge-Based Syst., vol. 317, no. May, p. 113370, May 2025, doi: 10.1016/j.knosys.2025.113370.
[31] D. Kiyasseh et al., “A vision transformer for decoding surgeon activity from surgical videos,” Nat. Biomed. Eng., vol. 7, no. 6, pp. 780–796, Mar. 2023, doi: 10.1038/s41551-023-01010-8.
[32] Y. Wang et al., “InternVideo2: Scaling Foundation Models for Multimodal Video Understanding,” in Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), Springer, Cham, 2025, pp. 396–416. doi: 10.1007/978-3-031-73013-9_23.
[33] J. Do and M. Kim, “SkateFormer: Skeletal-Temporal Transformer for Human Action Recognition,” in Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), vol. 15099 LNCS, Springer, Cham, 2025, pp. 401–420. doi: 10.1007/978-3-031-72940-9_23.
[34] I. Funke, D. Rivoir, S. Krell, and S. Speidel, “TUNeS: A Temporal U-Net With Self-Attention for Video-Based Surgical Phase Recognition,” IEEE Trans. Biomed. Eng., vol. 72, no. 7, pp. 2105–2119, Jul. 2025, doi: 10.1109/TBME.2025.3535228.
[35] A. Pérez, S. Rodríguez, N. Ayobi, N. Aparicio, E. Dessevres, and P. Arbeláez, “MuST: Multi-scale Transformers for Surgical Phase Recognition,” in Medical Image Computing and Computer Assisted Intervention – MICCAI 2024, Springer, Cham, 2024, pp. 422–432. doi: 10.1007/978-3-031-72089-5_40.
[36] X. Zhang et al., “SPRMamba: Surgical Phase Recognition for Endoscopic Submucosal Dissection with Mamba,” IEEE Sens. J., pp. 1–1, Jan. 2026, doi: 10.1109/JSEN.2025.3649151.
[37] C. I. Nwoye et al., “Rendezvous: Attention mechanisms for the recognition of surgical action triplets in endoscopic videos,” Med. Image Anal., vol. 78, no. May, p. 102433, May 2022, doi: 10.1016/j.media.2022.102433.
[38] O. Alabi, T. Vercauteren, and M. Shi, “Multitask learning in minimally invasive surgical vision: A review,” Med. Image Anal., vol. 101, no. April, p. 103480, Apr. 2025, doi: 10.1016/j.media.2025.103480.
[39] Y. Li, Z. Zhao, R. Li, and F. Li, “Deep learning for surgical workflow analysis: a survey of progresses, limitations, and trends,” Artif. Intell. Rev., vol. 57, no. 11, p. 291, Sep. 2024, doi: 10.1007/s10462-024-10929-6.
[40] C.-Y. Wang, A. Bochkovskiy, and H.-Y. M. Liao, “YOLOv7: Trainable Bag-of-Freebies Sets New State-of-the-Art for Real-Time Object Detectors,” in 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Jun. 2023, pp. 7464–7475. doi: 10.1109/CVPR52729.2023.00721.
[41] Z. Teed and J. Deng, “RAFT: Recurrent All-Pairs Field Transforms for Optical Flow,” in Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), vol. 12347 LNCS, Springer, Cham, 2020, pp. 402–419. doi: 10.1007/978-3-030-58536-5_24.
[42] A. Vaswani et al., “Attention is All you Need,” in Advances in Neural Information Processing Systems, 2017, p. 11. Accessed: Feb. 17, 2026. [Online]. Available: https://arxiv.org/abs/1706.03762.
[43] A. Dosovitskiy et al., “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,” in ICLR 2021 - 9th International Conference on Learning Representations, International Conference on Learning Representations, ICLR, Jun. 2021, p. 22. Accessed: Feb. 17, 2026. [Online]. Available: http://arxiv.org/abs/2010.11929.
[44] H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers & distillation through attention,” in Proceedings of Machine Learning Research, ML Research Press, Jan. 2021, pp. 10347–10357. Accessed: Feb. 17, 2026. [Online]. Available: http://arxiv.org/abs/2012.12877.
[45] I. Loshchilov and F. Hutter, “Decoupled Weight Decay Regularization,” in 7th International Conference on Learning Representations, ICLR 2019, International Conference on Learning Representations, ICLR, Jan. 2019, p. 19. Accessed: Feb. 17, 2026. [Online]. Available: http://arxiv.org/abs/1711.05101.

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
___________________________________________________________
International Journal of Advances in Intelligent Informatics
ISSN 2442-6571 (print) | 2548-3161 (online)
Organized by UAD and ASCEE Computer Society
Published by Universitas Ahmad Dahlan
W: http://ijain.org
E: info@ijain.org (paper handling issues)
andri.pranolo.id@ieee.org (publication issues)
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0

























Download