Reflections on Practical Paths of Empowering Embodied AI Development via Computer Vision in Robotic Systems
DOI:
https://doi.org/10.54097/tgds0y20Keywords:
Robotic system, computer vision, Embodied Artificial Intelligence, Perceptual Interaction, Intelligent empowermentAbstract
Currently, most commercial robots rely on fixed programs to perform repetitive tasks. The visual modules equipped on these devices have limited anti-interference capabilities, making it difficult to adapt to complex structures and variable environments, which in turn limits the practical effectiveness of robotic intelligence deployment. Drawing from hands-on experience in robot debugging and project implementation, this paper explores the integration of computer vision and embodied intelligence. Based on common application scenarios such as industrial sorting, power inspection, and intelligent services, it analyzes specific methods by which visual technologies assist robots in environmental perception, intelligent decision-making, motion adjustment, and model updating. The study identifies several common challenges encountered during industry deployment, including insufficient stability in recognizing complex scenes, difficulty balancing computational power with recognition accuracy, significant gaps between simulation training and real-world conditions, and a lack of unified application standards across industries. Addressing these practical issues, the paper proposes feasible improvement strategies from four perspectives: algorithm enhancement, computational power upgrade, simulation optimization, and the establishment of industry standards. The conclusions drawn from this research offer practical insights for advancing robot vision technology, expanding application scenarios, and promoting standardized development within the industry.
Downloads
References
[1] Lake, B. M., Ullman, T. D., Tenenbaum, J. B., et al. (2017). Building machines that learn and think like people. Behavioral and Brain Sciences, 40, e253.
[2] Howard, D., Eiben, A. E., Kennedy, D. F., et al. (2019). Evolving embodied intelligence from materials to machines. Nature Machine Intelligence, 1(1), 12–19.
[3] Jin, D., & Zhang, L. (2020). Embodied intelligence weaves a better future. Nature Machine Intelligence, 2(11), 663–664.
[4] Du, S., & Fei, Y. (2022). Embodied artificial intelligence: From perception to physical interaction. Nature Machine Intelligence, 4(11), 927–934.
[5] Cadena, C., Carlone, L., Carrillo, H., et al. (2016). Past, present, and future of simultaneous localization and mapping: toward the robust perception age. IEEE Transactions on Robotics, 32(6), 1309–1332.
[6] Lepetit, V., & Fua, P. (2015). Monocular 3D tracking of rigid and articulated objects. International Journal of Computer Vision, 111(2), 147–170.
[7] Bohg, A., Morales, A., Asfour, T., et al. (2014). Data driven grasp synthesis—a survey. IEEE Transactions on Robotics, 30(2), 289–309.
[8] Yang, S., & Scherer, S. (2019). Direct monocular odometry with active feature selection for robust navigation. IEEE Transactions on Robotics, 35(4), 963–979.
[9] Zhang, J., & Singh, S. (2018). LOAM: lidar odometry and mapping in real time. IEEE Transactions on Robotics, 34(4), 1056–1070.
[10] Ren, L., Dong, J., Liu, S., et al. (2025). Embodied intelligence toward future smart manufacturing in the era of AI foundation model. IEEE/ASME Transactions on Mechatronics, 30(4), 2632–2642.
[11] Hussein, A., Gaber, M. M., Elyan, E., et al. (2017). Imitation learning: a survey of learning methods. ACM Computing Surveys, 50(2), 1–35.
[12] Zhang, D. (2022). Explainable hierarchical imitation learning for robotic drink pouring. IEEE Transactions on Automation Science and Engineering, 19(4), 3871–3887.
[13] Levine, S., Finn, C., Darrell, T., et al. (2016). End to end training of deep visuomotor policies. Journal of Machine Learning Research, 17, 1–40.
[14] Pinto, L., & Gupta, A. (2017). Supersizing self supervision: learning to grasp from 50k tries and 700 robot hours. The International Journal of Robotics Research, 36(10 11), 1231–1250.
[15] Zhu, Y. (2020). Target driven visual navigation in indoor scenes via deep reinforcement learning. IEEE Transactions on Pattern Analysis and Machine Intelligence, 42(10), 2409–2423.
[16] Shridhar, M., Manuelli, L., & Fox, D. (2023). CLIPort: what and where pathways for robotic manipulation. The International Journal of Robotics Research, 42(1 2), 112–130.
[17] Liu, X., Li, Y., & Wang, H. (2026). Vision language action foundation models for robotic embodied intelligence: survey and open challenges. IEEE Transactions on Cybernetics, 56(2), 789–804.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Journal of Computer Science and Artificial Intelligence

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.








