Articles
Vol. 1 No. 4 (2026)
Causal Tracklet Graphs for Detecting Process Deviations in Partially Observed Factory Video
Department of Electrical Engineering, School of Engineering, Universidade Federal do Rio Grande do Sul, Porto Alegre, Rio Grande do Sul, Brazil
-
Submitted
-
August 11, 2026
-
Published
-
August 25, 2026
Abstract
The deployment of automated monitoring systems in modern manufacturing environments is heavily constrained by partial observability, a pervasive issue resulting from dynamic occlusions, limited camera fields of view, and complex spatial layouts of machinery. Traditional video anomaly detection methodologies rely heavily on associative learning paradigms, which frequently fail to distinguish spurious environmental correlations from genuine process deviations. This paper introduces Causal Tracklet Graphs, a novel computational framework specifically designed to detect process deviations in partially observed factory videos by modeling the underlying causal structure of sequential assembly tasks. We first extract temporal tracklets of workers, tools, and components, formulating them into a unified spatio-temporal graph representation. By integrating a structural causal model into the graph architecture, we perform counterfactual reasoning to mitigate the confounding effects of environmental noise, shifting illumination, and camera viewpoint biases. Through explicit causal interventions, specifically utilizing back-door adjustments, the proposed framework successfully isolates the true causal effects of localized actions on overarching process outcomes. Extensive experiments conducted on simulated and real-world large-scale industrial datasets demonstrate that the proposed method significantly outperforms current state-of-the-art anomaly detection models, particularly in scenarios characterized by severe occlusions and rare deviation patterns. The integration of rigorous causal inference with graph-based video representations offers a highly robust, interpretable, and scalable solution for quality assurance and workflow monitoring in advancing Industry 4.0 applications.
References
- 1. Zhang, W., Huang, M., Zhou, Y., Zhang, J., Yu, J., Wang, J., & Xu, L. (2024). Both2hands: Inferring 3d hands from both text prompts and body dynamics. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 2393-2404).
- 2. Cong, P., Xu, Y., Ren, Y., Zhang, J., Xu, L., Wang, J., ... & Ma, Y. (2023, June). Weakly supervised 3d multi-person pose estimation for large-scale scenes based on monocular camera and single lidar. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 37, No. 1, pp. 461-469).
- 3. Jing, G., Gao, P., Lee, Y., Hu, Y., & Zhang, H. (2025). 3D-aided pedestrian representation learning for video-based person re-identification. IEEE Transactions on Circuits and Systems for Video Technology, 35 (12), 12830-12845.
- 4. Lee, Y., Gao, P., Xu, Y., & Fan, W. (2025, October). How Do Optical Flow and Textual Prompts Collaborate to Assist in Audio-Visual Semantic Segmentation?. In 2025 IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 23342-23352). IEEE.
- 5. Su, Z., Wei, H., Cen, K., Wang, Y., Chen, G., Yuan, C., & Chu, X. (2026). Generation enhances understanding in unified multimodal models via multi-representation generation. arXiv preprint arXiv:2601.21406.
- 6. Qu, W., Shao, Y., Meng, L., Huang, X., & Xiao, L. (2024, June). A conditional denoising diffusion probabilistic model for point cloud upsampling. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 20786-20795). IEEE.
- 7. Li, X., Yang, F., Chen, L., & Cai, H. (2016, July). Saliency transfer: An example-based method for salient object detection. In IJCAI (pp. 3411-3417).
- 8. Qu, W., Wang, J., Gong, Y., Huang, X., & Xiao, L. (2025, June). An end-to-end robust point cloud semantic segmentation network with single-step conditional diffusion models. In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 27325-27335). IEEE.
- 9. Wang, Z., Yuan, R., Geng, Z., Li, H., Qu, X., Li, X., ... & Zhang, K. (2025, October). Singing timbre popularity assessment based on multimodal large foundation model. In Proceedings of the 33rd ACM International Conference on Multimedia (pp. 12227-12236).
- 10. Chen, L., Wang, J., Mortlock, T., Khargonekar, P., & Al Faruque, M. A. (2025, June). Hyperdimensional uncertainty quantification for multimodal uncertainty fusion in autonomous vehicles perception. In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 22306-22316). IEEE.
- 11. Xia, Y., Lu, Y., Gao, Y., & Ma, J. (2024, March). Locality preserving refinement for shape matching with functional maps. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 38, No. 6, pp. 6207-6215).
- 12. Xia, Y., & Ma, J. (2026). Locality Optimization Refinement with Deformation for Shape Matching via Functional Maps. International Journal of Computer Vision, 134(2), 76.
- 13. Yang, Y., Liao, Y. H., Bortins, I., Baldwin, D. P., & Zhang, S. (2024). Unidirectional structured light system calibration with auxiliary camera and projector.Optics and Lasers in Engineering,175, 107984.
- 14. Yang, Y., & Zhang, S. (2024). Pixelwise calibration method for a telecentric structured light system.Applied Optics,63(10), 2562-2569.
- 15. Wang, H., Xu, Q., Wang, C., Xue, T., Peng, C., Chen, W., & Lin, F. (2026). Bad Seeing or Bad Thinking? Rewarding Perception for Multimodal Reasoning.arXiv preprint arXiv:2605.14054.
- 16. Chen, L., Yang, Y., & Zhang, S. (2025). Auto-focusing and auto-exposure method for real-time structured-light 3D imaging.Optical Engineering,64(4), 044105-044105.
- 17. Wang, E., Fan, W., & Zhang, D. (2026). ADM-DP: Adaptive Dynamic Modality Diffusion Policy through Vision-Tactile-Graph Fusion for Multi-Agent Manipulation. arXiv preprint arXiv:2602.21622.
- 18. Wang, H., Wei, C., Ren, W., Liu, J., Lin, F., & Chen, W. (2026). Rationalrewards: Reasoning rewards scale visual generation both training and test time.arXiv preprint arXiv:2604.11626.
- 19. Li, Q., Liu, Z., Luo, W., Luo, T., & Hou, C. (2026). Correcting Visual Blur Induced by Attention Distraction to Reduce Hallucinations: Algorithm and Theory. arXiv preprint arXiv:2605.24602.
- 20. Li, Q., Luo, T., Jiang, M., Jiang, Z., Hou, C., & Li, F. (2025, April). Semi-supervised multi-view multi-label learning with view-specific transformer and enhanced pseudo-label. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 39, No. 17, pp. 18430-18438).
- 21. Du, Y., & Villarrubia-González, G. (2026). Lightweight Skin Lesion Segmentation for Edge Deployment: A Critical Review of Architectures, Accuracy Compensation, and Clinical Translation. IEEE Access.
- 22. Chen, L., Yang, Y., & Zhang, S. (2024, September). Rapid autofocusing method for digital fringe projection techniques. In Interferometry and Structured Light 2024(Vol. 13135, pp. 92-97). SPIE.
- 23. Ding, Y., Li, Z., Huang, D., Li, Z., & Zhang, K. (2022, October). Enhancing multi-view stereo with contrastive matching and weighted focal loss. In 2022 IEEE International Conference on Image Processing (ICIP) (pp. 821-825). IEEE.
- 24. Chang, C. C., Lu, T. C., & Chang, C. C. (2026). Subband-Guided Hybrid Multi-Axis Attention Network for Frequency-Aware Image Super-Resolution. Journal of Machine Learning Advances, 1(1), 1-23.
- 25. Li, Z., Jiang, L., Zhao, Y., Chen, Y. C., Wang, X., Chen, W., & Peng, Y. (2026). PatternGSL: A Structured Specification Language for Template-Free and Simulation-Ready 3D Garments. arXiv preprint arXiv:2606.24564.
- 26. Girshick, R. (2015). Fast R-CNN. In Proceedings of ICCV.
- 27. Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., et al. (2023). Segment anything. In Proceedings of ICCV.
- 28. Sermanet, P., Eigen, D., Zhang, X., Mathieu, M., Fergus, R., & LeCun, Y. (2014). OverFeat: Integrated recognition, localization and detection using convolutional networks. In Proceedings of ICLR.
- 29. Huhle, T., Torun, O., Song, W., & Breuil, C. (2026, June). Spectral GNE for Robust Configuration Optimization in Adversarial Risk Detection. In 2026 7th International Conference on Big Data & Artificial Intelligence & Software Engineering (ICBASE) (pp. 406-410). IEEE.
- 30. Yang, Y., Hu, H., Mao, Y., Zhang, J., Wu, C., Jiang, Y., ... & Zhang, C. (2026). OPRIDE: Offline Preference-based Reinforcement Learning via In-Dataset Exploration. arXiv preprint arXiv:2604.02349.