Real-time perception and control systemfor flexible handling with industrial robots
DOI:
https://doi.org/10.17979/ja-cea.2026.47.13705Keywords:
Perception and sensing, Robotics technology, Robot manipulators, Autonomous robotic systemsAbstract
This work presents the development of a flexible robotic manipulation system based on computer vision for industrial applications. The main goal is to identify and locate objects through instance segmentation and automatically compute their pose (position and orientation) to enable robotic-arm manipulation. The system follows a modular architecture built on the ROS 2 ecosystem, integrating visual perception components, manipulator control, and a graphical interface. A ZED Mini stereo camera captures RGB-D images, and a YOLO-based instance segmentation model identifies objects and generates corresponding masks. These, together with depth information, allow a robust grasping point to be computed. Finally, the proposed system incorporates a verification post-grasping mechanism. The results demonstrate robust and real-time performance, validating the application of advanced vision techniques in unstructured industrial environments. The code is publicly available in the repository (https://github.com/PedroMM03/semantic-segmentation-adaptive-manipulation).
References
Hartley, R., Zisserman, A., 2004. Multiple View Geometry in Computer Vision, 2nd Edition. Cambridge University Press.
Khanam, R., Hussain, M., 2024. YOLOv11: An overview of the key architectural enhancements. arXiv preprint arXiv:2408.00714.
Liu, Y., Graf, C., Spies, M., Keuper, M., 2025. Segment any repeated object. In: International Conference on Robotics and Automation. pp. 16042–16049. DOI: 10.1109/ICRA55743.2025.11128131
Redmon, J., Divvala, S., Girshick, R., Farhadi, A., 2016. You only look once: Unified, real-time object detection. In: IEEE Conference on Computer Visión and Pattern Recognition. pp. 779–788. DOI: 10.1109/CVPR.2016.91
ten Pas, A., Gualtieri, M., Saenko, K., Platt, R., 2017. Grasp pose detection in point clouds. The International Journal of Robotics Research 36 (13-14), 1455–1473. DOI: 10.1177/0278364917735594
Utomo, T. W., Cahyadi, A. I., Ardiyanto, I., 2021. Suction-based grasp point estimation in cluttered environment for robotic manipulator using Deep learning-based affordance map. International Journal of Automation and Computing 18 (2), 277–287. DOI: https://doi.org/10.1007/s11633-020-1260-1
Zapata-Impata, B. S., Gil, P., Pomares, J., Torres, F., 2019. Fast geometrybased computation of grasping points on three-dimensional point clouds. International Journal of Advanced Robotic Systems 16 (1), 1729881419831846. DOI: 10.1177/1729881419831846
Zhao, M., Zuo, G., Yu, S., Luo, Y., Liu, C., Gong, D., 2025. Language-guided category push–grasp synergy learning in clutter by efficiently perceiving object manipulation space. IEEE Transactions on Industrial Informatics 21 (2), 1783–1792. DOI: 10.1109/TII.2024.3488774
Zhong, X., Gong, T., Zhong, X., Liu, Q., Hu, H., 2025. Region-aware 6D grasping for industrial bin-picking: A Sim2Real label self-generation and hybrid evaluation framework. In: International Conference on Intelligent Robots and Systems. pp. 9448–9455. DOI: 10.1109/IROS60139.2025.11247129
Zhuang, C., Wang, Z., Zhao, H., Ding, H., 2021. Semantic part segmentation method based 3D object pose estimation with RGB-D images for binpicking. Robotics and Computer-Integrated Manufacturing 68, 102086. DOI: https://doi.org/10.1016/j.rcim.2020.102086
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Pedro Martín-Martos, Ricardo Vázquez-Martín

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.