Collaboration Efficiency with Gesture Recognition Interfaces in Mixed Reality Studios: Causal Modeling

Authors

  • Lerato D. Maseko Department Mechanical and Mechatronic Engineering, Faculty of Engineering, Stellenbosch University, Stellenbosch, Western Cape, South Africa Author

Keywords:

Mixed Reality, Gesture Recognition, Causal Modeling, Collaboration Efficiency, Human-Computer Interaction

Abstract

The rapid integration of mixed reality environments into professional and creative workflows has necessitated the development of more intuitive and efficient user interfaces. Among these, gesture recognition interfaces have emerged as a prominent solution, promising to reduce cognitive load and enhance spatial interaction. However, evaluating their true impact on collaborative efficiency remains challenging due to the presence of numerous confounding variables such as user experience, task complexity, and spatial cognitive abilities. Traditional observational studies often fail to isolate the causal effect of interface types on performance metrics, relying instead on correlational analyses that may lead to spurious conclusions. This paper introduces a comprehensive causal modeling framework to assess collaboration efficiency in mixed reality studios. By employing directed acyclic graphs and propensity score matching, this study systematically controls for environmental and user-specific confounders. The research design involves a controlled experimental setup where participants engage in collaborative spatial design tasks using either advanced gesture recognition interfaces or conventional handheld controllers. The empirical findings provide robust causal evidence indicating that while gesture interfaces significantly improve subjective feelings of co-presence and naturalistic interaction, their actual impact on objective task completion times is heavily moderated by the baseline spatial reasoning skills of the users. This study bridges a critical gap in human-computer interaction literature by transitioning from correlational observations to rigorous causal inference, offering actionable insights for the design and implementation of future mixed reality collaborative systems.

References

1. Farooq, M.A.; Corcoran, P.; Rotariu, C.; Shariff, W. Object detection in thermal spectrum for advanced driver-assistance systems (ADAS). IEEE Access 2021, 9, 156465–156481.

2. Shih-Cheng, H.; Anuj, P.; Malte, J.; Matthew, P.L.; Serena, Y.; Akshay, S.C. Self-supervised learning for medical image classification: A systematic review and implementation guidelines. npj Digit. Med. 2023, 6, 74.

3. Farahnakian, F.; Movahedi, P.; Poikonen, J.; Lehtonen, E.; Makris, D.; Heikkonen, J. Comparative analysis of image fusion methods in marine environment. In Proceedings of the 2019 IEEE International Symposium on Robotic and Sensors Environments (ROSE), Ottawa, ON, Canada, 17–18 June 2019; IEEE: New York, NY, USA, 2019.

4. Li, A.; Cheng, H.; Hu, S.; Liu, X.; Tang, J.; Lin, L. Learning collaborative sparse representation for grayscale-thermal tracking. IEEE Trans. Image Process. 2016, 25, 5743–5756.

5. Jaimes, A.; Sebe, N. Multimodal human–computer interaction: A survey. Comput. Vis. Image Underst. 2007, 108, 116–134.

6. Tran, K.; Nguyen, D.; Nguyen, D.; Tran, N.; Nguyen, H. Decision fusion from visible and infrared images for human emotion detection. Int. J. Adv. Eng. 2020, 3, 8–11.

7. Jiang, Y.; Li, W.; Hossain, M.S.; Chen, M.; Alelaiwi, A.; Al-Hammadi, M. A snapshot research and implementation of multimodal information fusion for data-driven emotion recognition. Inf. Fusion 2020, 53, 209–221.

8. Wang, F.; Chen, Y.; Wu, F.; Li, X. Textray: Contour-based geometric modeling for arbitrary-shaped scene text detection. In Proceedings of the 28th ACM International Conference on Multimedia, Virtual, 12–16 October 2020; pp. 111–119.

9. Brightman, N.; Fan, L. A brief overview of the current state, challenging issues and future directions of point cloud registration. ISPRS Ann. Photogramm. Remote Sens. Spat. Inf. Sci. 2022, 10, 17–23.

10. Ramesh, M.; Chahal, D.; Singhal, R. Multicloud Deployment of AI Workflows Using FaaS and Storage Services. In 15th International Conference on COMmunication Systems & NETworkS (COMSNETS); IEEE: Piscataway, NJ, USA, 2023; pp. 269–277.

11. Anitha, M.; Keerthana, R.; Regunath, E.; Hariram, R. Smart Thyroid: A Deep Learning-Driven SVM Model for Thyroid Diagnosis. In Proceedings of the IEEE International Conference on Intelligent Sustainable Systems (ICISS), Tirunelveli, India, 12–14 March 2025; pp. 1633–1639.

12. Qian, T.; Zhou, Y.; Yao, J.; Ni, C.; Asif, S.; Chen, C.; Lv, L.; Ou, D.; Xu, D. Deep learning based analysis of dynamic video ultrasonography for predicting cervical lymph node metastasis in papillary thyroid carcinoma. Endocrine 2025, 87, 1060–1069.

13. Osin, V.; Cichocki, A.; Burnaev, E. Fast multispectral deep fusion networks. Bull. Pol. Acad. Sci. Tech. Sci. 2018, 66, 875–889.

14. Tang, C.; Tian, G.Y.; Wu, J. Segmentation-oriented compressed sensing for efficient impact damage detection on CFRP materials. IEEE/ASME Trans. Mechatronics 2020, 26, 2528–2537.

15. Li, C.; Liang, X.; Lu, Y.; Zhao, N.; Tang, J. RGB-T object tracking: Benchmark and baseline. Pattern Recognit. 2019, 96, 106977.

16. Torabi, A.; Massé, G.; Bilodeau, G.-A. An iterative integrated framework for thermal–visible image registration, sensor fusion, and people tracking for video surveillance applications. Comput. Vis. Image Underst. 2012, 116, 210–221.

17. Yang, F.; Cheng, I. Visible-infrared features fusion based object detection. In Proceeding of the 2023 IEEE International Conference on Systems, Man, and Cybernetics (SMC), Honolulu, Oahu, HI, USA, 1–4 October 2023; IEEE: New York, NY, USA, 2023.

18. Siddiqui, I.F.H.; Javaid, A.Y. A multimodal facial emotion recognition framework through the fusion of speech with visible and infrared images. Multimodal Technol. Interact. 2020, 4, 46.

19. Zhang, C.; Yu, M.; Wang, W.; Yan, F. Enabling Cost-Effective, SLO-Aware Machine Learning Inference Serving on Public Cloud. IEEE Trans. Cloud Comput. 2022, 10, 1765–1779.

20. Siddiqui, I.F.H.; Dhakal, P.; Yang, X.; Javaid, A.Y. A survey on databases for multimodal emotion recognition and an introduction to the VIRI (visible and InfraRed image) database. Multimodal Technol. Interact. 2022, 6, 47.

21. Roseman, J. Emotional behaviors, emotivational goals, emotion strategies: Multiple levels of organization integrate variable and consistent responses. Emot. Rev. 2011, 3, 434–443.

22. Shao, J.; Huang, X.; Gao, T.; Cao, J.; Wang, Y.; Zhang, Q.; Lou, L.; Ye, J. Deep learning-based image analysis of eyelid morphology in thyroid-associated ophthalmopathy. Quant. Imaging Med. Surg. 2023, 13, 1592–1604.

23. Kant, V.; Jagadish, S. Thyroid Disease Detection Through EfficientNetB2: Enhancing Early Diagnosis with Deep Learning. In Proceedings of the OPJU International Technology Conference (OTCON) on Smart Computing for Innovation and Advancement in Industry 5.0, Raigarh, India, 9–11 April 2025; pp. 1–6.

24. NVIDIA Corporation. System Management Interface (SMI). Available online: https://developer.nvidia.com/system-management-interface (accessed on 30 March 2026).

25. Li, Q.; Han, G.; Liu, P.; Yang, H.; Luo, H.; Wu, J. An infrared-visible image registration method based on the constrained point feature. Sensors 2021, 21, 1188.

26. Hou, Z.; Yang, C.; Sun, Y.; Ma, S.; Yang, X.; Fan, J. An object detection algorithm based on infrared-visible dual modal feature fusion. Infrared Phys. Technol. 2024, 137, 105107.

27. Nathani, A.; Chaudhary, S.; Somani, G. Policy based resource allocation in IaaS cloud. Future Gener. Comput. Syst. 2012, 28, 94–103.

28. Docker Inc. Docker: Accelerated Container Application Development. Available online: https://www.docker.com/ (accessed on 30 March 2026).

Downloads

Published

2026-03-29

Issue

Section

Articles