Brain-Inspired Generative AI Revives Legacy Industrial Robots Toward Sustainable Automation

Zhihao Liu , Tianyu Wang , Ruirui Zhong , Xi Vincent Wang , Mian Li , Lihui Wang

Engineering ›› : 202605020

PDF (7662KB)
Engineering ›› :202605020 DOI: 10.1016/j.eng.2026.05.020
research-article
Brain-Inspired Generative AI Revives Legacy Industrial Robots Toward Sustainable Automation
Author information +
History +
PDF (7662KB)

Abstract

In recent decades, industrial robots have been widely used in automation. However, their reliance on specific process-oriented programs with various programming languages limits their adaptability. Moreover, advancements in embodied intelligence enabled a range of unicorns with innovative humanoid and quadruped robots to work in factories. From the perspective of sustainability, this paper is motivated by augmenting existing industrial robots already powerful in accuracy and payload. In this paper, PATON, a universal and modularized design inspired by the biological scheme of the human brain and the von Neumann architecture for modern computers, is proposed. PATON accepts natural-language commands and raw environmental images; performs scene understanding, semantic task reasoning and planning; and executes programming-free motion planning and control. PATON, which is implemented and validated on two ABB industrial robots in an automation cell, enables legacy industrial robots to conduct tasks with advanced machine intelligence and improves sustainability by revitalizing legacy equipment.

Keywords

Generative artificial intelligence / Industrial robots / Industrial automation / Smart manufacturing

Cite this article

Download citation ▾
Zhihao Liu, Tianyu Wang, Ruirui Zhong, Xi Vincent Wang, Mian Li, Lihui Wang. Brain-Inspired Generative AI Revives Legacy Industrial Robots Toward Sustainable Automation. Engineering 202605020 DOI:10.1016/j.eng.2026.05.020

登录浏览全文

4963

注册一个新账户 忘记密码

References

[1]

Gupta A, Savarese S, Ganguli S, Fei—Fei L . Embodied intelligence via learning and evolution. Nat Commun 2021; 12(1):5721.

[2]

Epstein Z, Hertzmann A ; The Investigators of Human Creativity. Art and the science of generative AI. Science 2023; 380:1110-1.

[3]

Malik AA, Masood T, Brem A . Intelligent humanoid robots in manufacturing. In: Companion Proceedings of the 2024 ACM/IEEE International Conference on Human—Robot Interaction; 2024 Mar 11—15; Boulder, CO, USA. New York City: Association for Computing Machinery; 2024. p. 20-7.

[4]

Mon—Williams R, Li G, Long R, Du W, Lucas CG . Embodied large language models enable robots to complete complex tasks in unpredictable environments. Nat Mach Intell 2025; 7(4):592-601.

[5]

Wang P, Zhang LY, Tzachor A, Chen WQ . E—waste challenges of generative artificial intelligence. Nat Comput Sci 2024; 4(11):818—23.

[6]

Fan H, Liu X, Fuh JYH, Lu WF, Li B . Embodied intelligence in manufacturing: leveraging large language models for autonomous industrial robotics. J Intell Manuf 2025; 36(2):1141—57.

[7]

OpenAI; Achiam J, Adler S, Agarwal S, Ahmad L, Akkaya I, et al. GPT—4 technical report. 2023. arXiv:2303.08774.

[8]

Guo D, Yang D, Zhang H, Song J, Wang P, Zhu Q, et al. DeepSeek—R1 incentivizes reasoning in LLMs through reinforcement learning. Nature 2025; 645(8081):633—8.

[9]

Driess D, Xia F, Sajjadi MSM, Lynch C, Chowdhery A, Ichter B, et al. PaLM—E: an embodied multimodal language model. In: Proceedings of the 40th International Conference on Machine Learning; 2023 Jul 23—29; Honolulu, HI, USA. Brookline: JMLR.org; 2023. p. 8469-88.

[10]

Han K, Wang Y, Chen H, Chen X, Guo J, Liu Z, et al. A survey on vision transformer. IEEE Trans Pattern Anal Mach Intell 2023; 45(1):87-110.

[11]

Zitkovich B, Yu T, Xu S, Xu P, Xiao T, Xia F, et al. RT—2: vision—language—action models transfer web knowledge to robotic control. In: Proceedings of the 7th Conference on Robot Learning; 2023 Nov 6—9; Atlanta, GA, USA. New York City: PMLR; 2023. p. 2165-83.

[12]

Kim MJ, Pertsch K, Karamcheti S, Xiao T, Balakrishna A, Nair S, et al. OpenVLA: an open—source vision—language—action model. 2024. arXiv:2406.09246.

[13]

Pertsch K, Stachowicz K, Ichter B, Driess D, Nair S, Vuong Q, et al. Fast: efficient action tokenization for vision—language—action models. 2025. arXiv:2501.09747.

[14]

Zhao TZ, Kumar V, Levine S, Finn C . Learning fine—grained bimanual manipulation with low—cost hardware. 2023. arXiv:2304.13705.

[15]

Intelligence P, Black K, Brown N, Darpinian J, Dhabalia K, Driess D, et al. 𝜋0.5: a vision—language—action model with open—world generalization. 2025. arXiv:2504.16054.

[16]

Black K, Brown N, Driess D, Esmail A, Equi M, Finn C, et al. 𝜋0: a vision—language—action flow model for general robot control. 2024. rXiv:2410.24164.

[17]

Liu S, Wu L, Li B, Tan H, Chen H, Wang Z, et al. RDT—1B: a diffusion foundation model for bimanual manipulation. 2024. arXiv:2410.07864.

[18]

Liu J, Chen H, An P, Liu Z, Zhang R, Gu C, et al. HybridVLA: collaborative diffusion and autoregression in a unified vision—language—action model. 2025. arXiv:2503.10631.

[19]

O’Neill A, Rehman A, Maddukuri A, Gupta A, Padalkar A, Lee A, et al. Open X—embodiment: robotic learning datasets and RT—X models: Open X—embodiment collaboration0 . In: Proceedings of the 2024 IEEE International Conference on Robotics and Automation (ICRA); 2024 May 13—17; Yokohama, Japan. Piscataway: IEEE; 2024. p. 6892-903.

[20]

Tamizi MG, Yaghoubi M, Najjaran H . A review of recent trend in motion planning of industrial robots. Int J Intell Robot Appl 2023; 7(2):253—74.

[21]

International Federation of Robotics. World robotics 2024: Industrial robots. Report. Frankfurt: International Federation of Robotics ; 2024.

[22]

Bilancia P, Schmidt J, Raffaeli R, Peruzzini M, Pellicciari M . An overview of industrial robots control and programming approaches. Appl Sci 2023; 13(4):2582.

[23]

IFR. World Robotics 2024 — Industrial Robots [Internet]. Frankfurt: IFR;2024 [cited 2025 Dec 09]. Available from:https://ifr.org/img/worldrobotics/Executive_Summary_WR_2024_Industrial_Robots.pdf.

[24]

Awais M, Naseer M, Khan S, Anwer RM, Cholakkal H, Shah M, et al. Foundation models defining a new era in vision: a survey and outlook. IEEE Trans Pattern Anal Mach Intell 2025; 47(4):2245-64.

[25]

Shanahan M. Talking about large language models. Commun ACM 2024; 67(2):68-79.

[26]

Wang S, Tan Z, Wang Y, Zhou Z, Zhu D . A novel paradigm of robotic machining towards embodied intelligent manufacturing: case study on paint defect repair. Robot Comput—Integr Manuf 2026; 98:103184.

[27]

Liu Y, Wang Q, Zhu Z, Wang Z, Lai M, Li R, et al. Embodied artificial intelligent industrial robot: system architecture, key technologies and case study. Comput Integr Manuf Syst 2025; 31(12):4513-41. Chinese.

[28]

von Neumann J . First draft of a report on the EDVAC. IEEE Ann Hist Comput 1993; 15(4):27-75.

[29]

Nazari K, Mandil W, Santello M, Park S, Ghalamzan—E A . Bioinspired trajectory modulation for effective slip control in robot manipulation. Nat Mach Intell 2025;7(7):1119-28.

[30]

Merel J, Botvinick M, Wayne G . Hierarchical motor control in mammals and machines. Nat Commun 2019; 10(1):5489.

[31]

Liu H, Li C, Wu Q, Lee YJ . Visual instruction tuning. In: Proceedings of the 37th International Conference on Neural Information Processing Systems; 2023 Dec 10—16; New Orleans, LA, USA. Red Hook: Curran Associates Inc.; 2023. p. 34892—916.

[32]

Li J, Li D, Savarese S, Hoi S . BLIP—2: bootstrapping language—image pre—training with frozen image encoders and large language models. In: Proceedings of the 40th International Conference on Machine Learning; 2023 Jul 23—29; Honolulu, HI, USA. Brookline: JMLR.org;2023. p. 19730—42.

[33]

Asai A, Wu Z, Wang Y, Sil A, Hajishirzi H . Self—RAG: learning to retrieve, generate, and critique through self—reflection. In: Proceedings of the Twelfth International Conference on Learning Representations; 2024 May 7—11; Vienna, Austria. Washington, DC: ICLR; 2024. p. 9112-41.

[34]

Gu Q, Kuwajerwala A, Morin S, Jatavallabhula KM, Sen B, Agarwal A, et al. ConceptGraphs: Open—vocabulary 3D scene graphs for perception and planning. In: Proceedings of the 2024 IEEE International Conference on Robotics and Automation (ICRA); 2024 May 13—17; Yokohama, Japan. Piscataway: IEEE; 2024. p. 5021—8.

[35]

Liu S, Zeng Z, Ren T, Li F, Zhang H, Yang J, et al. Grounding DINO: marrying DINO with grounded pre—training for open—set object detection. In: Leonardis A, Ricci E, Roth S, Russakovsky O, Sattler T, Varol G, editors. Computer Vision— ECCV 2024. Cham: Springer; 2024. p. 38-55.

[36]

Ravi N, Gabeur V, Hu YT, Hu R, Ryali C, Ma T, et al. SAM 2: segment anything in images and videos. 2024. arXiv:2408.00714.

[37]

Wen B, Yang W, Kautz J, Birchfield S . FoundationPose: unified 6D pose estimation and tracking of novel objects. In: Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2024 Jun 16—22; Seattle, WA, USA. Piscataway: IEEE; 2024. p. 17868—79.

[38]

Yao S, Zhao J, Yu D, Du N, Shafran I, Narasimhan K, et al. ReAct: synergizing reasoning and acting in language models. In: Proceedings of the Eleventh International Conference on Learning Representations; 2023 May 1—5; Kigali, Rwanda. Washington, DC: ICLR; 2023. p. 1-33.

[39]

Huang W, Xia F, Xiao T, Chan H, Liang J, Florence P, et al. Inner Monologue: embodied reasoning through planning with language models. In: Proceedings of the 6th Conference on Robot Learning; 2022 Dec 14—18; Auckland, New Zealand. New York City: PMLR; 2023. p. 1769—82.

[40]

Shinn N, Cassano F, Gopinath A, Narasimhan K, Yao S . Reflexion: language agents with verbal reinforcement learning. In: Proceedings of the 37th International Conference on Neural Information Processing Systems; 2023 Dec 10—16; New Orleans, LA, USA. Red Hook: Curran Associates Inc.; 2023. p. 8634—52.

[41]

Sucan IA, Moll M, Kavraki LE . The open motion planning library. IEEE Robot Autom Mag 2012; 19(4):72-82.

[42]

Beeson P, Ames B . TRAC—IK: An open—source library for improved solving of generic inverse kinematics. In: Proceedings of the 2015 IEEE—RAS 15th International Conference on Humanoid Robots (Humanoids); 2015 Nov 3—5; Seoul, Republic of Korean. Piscataway: IEEE; 2015. p. 928—35.

[43]

Cohn T, Shaw S, Simchowitz M, Tedrake R . Constrained bimanual planning with analytic inverse kinematics. In: Proceedings of the 2024 IEEE International Conference on Robotics and Automation (ICRA); 2024 May 13—17; Yokohama, Japan. Piscataway: IEEE; 2024. p. 6935—42.

[44]

Degai K. Traditional robot programming vs AI and machine vision. Report. Frankfurt: IFR.

[45]

Macenski S, Foote T, Gerkey B, Lalancette C, Woodall W . Robot Operating System 2: design, architecture, and uses in the wild. Sci Robot 2022; 7(66):eabm6074.

[46]

Tao F, Qi Q . Make more digital twins. Nature 2019; 573(7775):490-1.

[47]

Tao F, Zhang H, Zhang C . Advancements and challenges of digital twins in industry. Nat Comput Sci 2024; 4(3):169-77.

[48]

Blankemeyer S, Wendorff D, Raatz A . Robotic evaluation framework for 6D object pose estimation accuracy. Procedia CIRP 2025; 134:1113—8.

[49]

Liu Z, Liu X, Cao Z, Gong X, Tan M, Yu J . High precision calibration for three—dimensional vision—guided robot system. IEEE Trans Ind Electron 2023; 70(1):624—34.

PDF (7662KB)

0

Accesses

0

Citation

Detail

Sections
Recommended

/