Vision–Language–Action Models for Agricultural Robotics: A Review Towards Embodied Intelligence

Xiaopei Yang , Jie He , Lei Ye , Zhaojie Wu , Yunchao Tang , Xiwen Luo

Engineering ›› : 202606021

PDF (8445KB)
Engineering ›› :202606021 DOI: 10.1016/j.eng.2026.06.021
research-article
Vision–Language–Action Models for Agricultural Robotics: A Review Towards Embodied Intelligence
Author information +
History +
PDF (8445KB)

Abstract

The shift toward Agriculture 5.0 demands robotic systems that can reason and act reliably in unstructured, high-entropy environments. Conventional agricultural robots, limited by rigid modular pipelines, often fail to generalize across diverse crop morphologies and rapidly changing field conditions. In contrast to existing surveys on general-purpose robotics, this paper provides a comprehensive review of vision–language–action (VLA) collaborative intelligence systems. It proposes a transformative framework that integrates multimodal foundation models to bridge the cognitive gap in agricultural robotics. We systematically examine VLA architectures along three core dimensions tailored to agricultural

Keywords

Agricultural robotics / Vision–language–action models / Embodied intelligence / Large language models / Precision agriculture

Cite this article

Download citation ▾
Xiaopei Yang, Jie He, Lei Ye, Zhaojie Wu, Yunchao Tang, Xiwen Luo. Vision–Language–Action Models for Agricultural Robotics: A Review Towards Embodied Intelligence. Engineering 202606021 DOI:10.1016/j.eng.2026.06.021

登录浏览全文

4963

注册一个新账户 忘记密码

References

[1]

Food and Agriculture Organization (FAO) of the United Nations. The future of food and agriculture drivers and triggers for transformation. Rome: Food and Agriculture Organization (FAO) of the United Nations; 2022.

[2]

Duckett T, Pearson S, Blackmore S, Grieve B, Wilson P . Agricultural robotics: the future of robotic agriculture (UK RAS white papers). Edinburgh: UK RAS Network; 2018.

[3]

Fountas S, Espejo García B, Kasimati A, Mylonas N, Darra N . The future of digital agriculture: technologies and opportunities. IT Prof 2020; 22(1): 24-8.

[4]

Oliveira LF, Moreira AP, Silva MF . Advances in agriculture robotics: a state—of—the—art review and challenges ahead. Robotics 2021; 10(2): 52.

[5]

Tang Y, Chen M, Wang C, Luo L, Li J, Lian G, et al. Recognition and localization methods for vision based fruit picking robots: a review. Front Plant Sci 2020; 11: 510.

[6]

Bhat VS, Wang Y . Revisiting the control systems of autonomous vehicles in the agricultural sector: a systematic literature review. IEEE Access 2025; 13: 21204—15.

[7]

Bechar A, Vigneault C . Agricultural robots for field operations: concepts and components. Biosyst Eng 2016; 149: 94-111.

[8]

Zhou H, Wang X, Au W, Kang H, Chen C . Intelligent robots for fruit harvesting: recent developments and future challenges. Precis Agric 2022; 23(5): 1856-907.

[9]

Zhai Z, Martínez JF, Beltran V, Martínez NL . Decision support systems for agriculture 4.0: survey and challenges. Comput Electron Agric 2020; 170: 105256.

[10]

Kamilaris A, Prenafeta Boldú FX . Deep learning in agriculture: a survey. Comput Electron Agric 2018; 147: 70-90.

[11]

Voulodimos A, Doulamis N, Doulamis A, Protopapadakis E . Deep learning for computer vision: a brief review. Comput Intell Neurosci 2018: 7068349.

[12]

Kroemer O, Niekum S, Konidaris G . A review of robot learning for manipulation: challenges, representations, and algorithms. J Mach Learn Res 2021; 22: 1-82.

[13]

Xu J, Sun Q, Han QL, Tang Y . When embodied AI meets Industry 5.0: human centered smart manufacturing. IEEE/CAA J Autom Sin 2025; 12: 485-501.

[14]

Saiz Rubio V, Rovira Más F . From smart farming towards agriculture 5.0: a review on crop data management. Agronomy 2020; 10(2): 207.

[15]

Vougioukas SG . Agricultural robotics. Annu Rev Control Robot Auton Syst 2019; 2(1): 365-92.

[16]

Firoozi R, Tucker J, Tian S, Majumdar A, Sun J, Liu W, et al. Foundation models in robotics: applications, challenges, and the future. Int J Robot Res 2025; 44(5): 701-39.

[17]

Liu Y, Chen W, Bai Y, Liang X, Li G, Gao W, et al. Aligning cyber space with physical world: a comprehensive survey on embodied AI. IEEE/ASME Trans Mechatron 2025; 30(6): 7253-74.

[18]

Brohan A, Brown N, Carbajal J, Chebotar Y, Dabis J, Finn C, et al. RT 1: robotics transformer for real world control at scale. In: Proceedings of Robotics: Science and Systems XIX; 2023 Jul 10—14; Daegu, Republic of Korea. Online: Robotics Science and Systems Online Proceedings; 2023.

[19]

Driess D, Xia F, Sajjadi MS, Lynch C, Chowdhery A, Wahid A, et al. PaLM E: an embodied multimodal language model. In: Agrawal P, Kroemer O, Burgard W, editors. Proceedings of the 40th International Conference on Machine Learning; 2023 Jul 23—29; Honolulu, HI, USA. London: PMLR; 2023. p. 8469-88.

[20]

Ghosh D, Walke HR, Pertsch K, Black K, Mees O, Dasari S, et al. Octo: an open source generalist robot policy. In: Proceedings of Robotics: Science and Systems XX; 2024 Jul 15—19; Delft, The Netherlands. Online: Robotics Science and Systems Online Proceedings; 2024.

[21]

Kim MJ, Pertsch K, Karamcheti S, Xiao T, Balakrishna A, Nair S, et al. OpenVLA: an open source vision—language—action model. In: Agrawal P, Kroemer O, Burgard W, editors. Proceedings of the 8th Conference on Robot Learning; 2025 Nov 6—9; Seoul, Republic of Korea. London: PMLR; 2025. p. 2679-713.

[22]

Vuong Q, Levine S, Walke HR, Pertsch K, Singh A, Doshi R, et al. Open X embodiment: robotic learning datasets and RT X models. In: Proceedings of the 2025 IEEE International Conference on Robotics and Automation; 2025 May 19—23; Atlanta, GA, USA. Piscataway: IEEE; 2025. p. 6892-903.

[23]

Kawaharazuka K, Oh J, Yamada J, Posner I, Zhu Y . Vision—language—action models for robotics: a review towards real—world applications. IEEE Access 2025; 13: 162467-504.

[24]

Gao B, Liu Y, Li Y, Li H, Li M, He W . A vision—language model for predicting potential distribution land of soybean double cropping. Front Environ Sci 2025; 12: 1515752.

[25]

Jha K, Doshi A, Patel P, Shah M . A comprehensive review on automation in agriculture using artificial intelligence. Artif Intell Agric 2019; 2: 1-12.

[26]

Hassanin M, Khan S, Tahtali M . Visual affordance and function understanding: a survey. ACM Comput Surv 2022; 54(3): 1-35.

[27]

Tian H, Wang T, Liu Y, Qiao X, Li Y . Computer vision technology in agricultural automation—a review. Inf Process Agric 2020; 7(1): 1-19.

[28]

Shamshiri R, Weltzien C, Hameed IAJ, Yule IE, Grift T, Balasundram SK, et al. Research and development in agricultural robotics: a perspective of digital farming. Int J Agric Biol Eng 2018; 11: 1-14.

[29]

Alayrac JB, Donahue J, Luc P, Miech A, Barr I, Laptev I, et al. Flamingo: a visual language model for few—shot learning. In: Advances in Neural Information Processing Systems; 2022 Nov 28—Dec 1 New Orleans, LA, USA. Red Hook: Curran Associates; 2022. p. 23716—36.

[30]

Radford A, Kim JW, Hallacy C, Ramesh A, Goh G, Agarwal S, et al. Learning transferable visual models from natural language supervision. In: Meila M, Zhang T, editors. Proceedings of the 38th International Conference on Machine Learning; 2021 Jul 18—24; Vienna, Austria. London: PMLR; 2021. p. 8748—63.

[31]

Li LH, Zhang P, Zhang H, Yang J, Li C, Yang J, et al. Grounded language image pre training. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; 2022 Jun 19—24; New Orleans, LA, USA. Piscataway: IEEE; 2022. p. 10965—75.

[32]

Minderer M, Gritsenko A, Stone A, Neumann M, Weissenborn D, Dosovitskiy A, et al. Simple open vocabulary object detection with vision transformers. In: Avidan S, Brostow G, Cissé M, Farinella GM, Hassner T, editors. Computer vision ECCV 2022; 2022 Aug 23—27; Tel Aviv, Israel. Cham: Springer; 2022. p. 728-55.

[33]

Jiang X, Wang J, Xie K, Cui C, Du A, Shi X, et al. PlantCaFo: an efficient few—shot plant disease recognition method based on foundation models. Plant Phenomics 2025; 7(1): 100024.

[34]

Zhu H, Qin S, Su M, Lin C, Li A, Gao J . Harnessing large vision and language models in agriculture: a review. Front Plant Sci 2025; 16: 1579355.

[35]

Ren T, Liu S, Zeng A, Lin J, Li K, Cao H, et al. Grounded SAM: assembling open—world models for diverse visual tasks. 2024. arXiv:2404.03875.

[36]

Lin B, Nie Y, Wei Z, Chen J, Ma S, Han J, et al. Navcot: boosting LLM based vision—and—language navigation via learning disentangled reasoning. IEEE Trans Pattern Anal Mach Intell 2025; 47(7): 5945-57.

[37]

Lu W, Wu X, Gao S, He W, Zhao Q, Zhang L, et al. Research on task decomposition and motion trajectory optimization of robotic arm based on VLA large model. In: Proceedings of the 2024 5th International Conference on Machine Learning and Computer Application; 2024 Oct 18—20; Hangzhou, China. Piscataway: IEEE; 2024. p. 80—5.

[38]

Liu H, Zhu Y, Kato K, Tsukahara A, Kondo I, Aoyama T, et al. Enhancing the LLM based robot manipulation through human—robot collaboration. IEEE Robot Autom Lett 2024; 9(8): 6904—11.

[39]

Wang J, Shi E, Hu H, Ma C, Liu Y, Wang X, et al. Large language models for robotics: opportunities, challenges, and perspectives. J Auton Intell 2025; 4(1): 52-64.

[40]

Liang J, Huang W, Xia F, Xu P, Hausman K, Ichter B, et al. Code as policies: language model programs for embodied control. In: Proceedings of the 2023 IEEE International Conference on Robotics and Automation; 2023 May 29—Jun 2; London, UK. Piscataway: IEEE; 2023. p. 9493—500.

[41]

Bian S, Zhang Y, Tian G, Miao Z, Wu EQ, Yang SX, et al. Large language model based task planning for service robots: a review. Biomim Intell Robot 2026; 6: 100274.

[42]

Shah D, Osiński B, Levine S . LM NAV: robotic navigation with large pre trained models of language, vision, and action. In: Agrawal P, Kroemer O, Burgard W, editors. Proceedings of the Conference on Robot Learning; 2023 Nov 6—9; Atlanta, GA, USA. PMLR; 2023. p. 492-504.

[43]

Zitkovich B, Yu T, Xu S, Xu P, Xiao T, Xia F, et al. RT 2: vision—language—action models transfer web knowledge to robotic control. In: Agrawal P, Kroemer O, Burgard W, editors. Proceedings of the Conference on Robot Learning; 2023 Nov 06—09; Atlanta, GA, USA. London: PMLR; 2023. p. 2165-83.

[44]

Chi C, Xu Z, Feng S, Cousineau E, Du Y, Burchfiel B, et al. Diffusion policy: visuomotor policy learning via action diffusion. Int J Robot Res 2024; 43(11—12): 1455-82.

[45]

Zhang J, Guo Y, Chen X, Wang YJ, Hu Y, Shi C, et al. HiRT: enhancing robotic control with hierarchical robot transformers. In: Agrawal P, Kroemer O, Burgard W, editors. Proceedings of the Conference on Robot Learning; 2025 Nov 6—9; Seoul, Republic of Korea. London: PMLR; 2025. p. 933—46.

[46]

Ichter B, Brohan A, Chebotar Y, Finn C, Hausman K, Herzog A, et al. Do as I can, not as I say: grounding language in robotic affordances. In: Liu K, Kulic D, Ichnowski J, editors. Proceedings of the 6th Conference on Robot Learning; 2022 Dec 14—18; Auckland, New Zealand. PMLR; 2023. p. 287-318.

[47]

Sapkota R, Cao Y, Roumeliotis KI, Karkee M . Vision—language—action models: concepts, progress, applications and challenges. 2025. arXiv:2505.04769.

[48]

Soncini N, Cremona J, Vidal E, García M, Castro G, Pire T . The Rosario dataset v2: multi modal dataset for agricultural robotics. Int J Robot Res 2025; 44(12): 1287-303.

[49]

Dolgopolyi R, Tsevas A . Bridging perception, language, and action: a survey and bibliometric analysis of VLM & VLA systems [Internet]. Durham: Research Square; 2025 Oct 26 [cited 2026 Aug 8]. Available from: https://www.researchsquare.com/article/rs—7935378/v1.

[50]

Wang J, Chen B, Li Y, Kang B, Chen Y, Tian Z . Declip: decoupled learning for open vocabulary dense perception. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; 2025 Jun 10—14; Nashville, TN, USA. Piscataway: IEEE; 2025. p. 14824—34.

[51]

Xu M, Cai D, Yin W, Wang S, Jin X, Liu X . Resource efficient algorithms and systems of foundation models: a survey. ACM Comput Surv 2025; 57(5): 1-39.

[52]

Li J, Li C, Zeng S, Tang Y, Yang C . A lightweight pineapple detection network based on YOLOv7—tiny for agricultural robot system. Comput Electron Agric 2025; 231: 109944.

[53]

Zhou Y, Yan H, Ding K, Cai T, Zhang Y . Few—shot image classification of crop diseases based on vision—language models. Sensors 2024; 24(18): 6109.

[54]

Kirillov A, Mintun E, Ravi N, Mao H, Rolland C, Gustafson L, et al. Segment anything. In: Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision; 2023 Oct 2—6; Paris, France. Piscataway: IEEE; 2023. p. 4015—26.

[55]

Li Y, Wang D, Yuan C, Li H, Hu J . Enhancing agricultural image segmentation with an agricultural segment anything model adapter. Sensors 2023; 23(18): 7884.

[56]

Ji W, Li J, Bi Q, Liu T, Li W, Cheng L . Segment anything is not always perfect: an investigation of SAM on different real—world applications. Mach Vis Appl 2024; 35(4): 617.

[57]

Huang Z, Wang H, Liu Y, Yang X, Wang Z, Liu X, et al. Segment anything model combined with multi—scale segmentation for extracting complex cultivated land parcels in high—resolution remote sensing images. Remote Sens 2024; 16(18): 3489.

[58]

Xiao X, Jiang Y, Wang Y . Key technologies for machine vision for picking robots: review and benchmarking. Mach Intell Res 2025; 22(1): 2-16.

[59]

Kerbl B, Kopanas G, Leimkühler T, Drettakis G . Gaussian splatting for real—time radiance field rendering. ACM Trans Graph 2023; 42(4): 1-14.

[60]

Gao J, Zhu C, Hu J, Fei D, Xu Z, Wang X . Seed 3D phenotyping across multiple crops using 3D Gaussian splatting. Agriculture 2025; 15(22): 2329.

[61]

Ku K, Jubery TZ, Rodriguez E, Balu A, Sarkar S, Krishnamurthy A, et al. SC—NeRF: NeRF—based point cloud reconstruction using a stationary camera for agricultural applications. In: Proceedings of the 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops; 2025 Jun 10—14; Nashville, TN, USA. Piscataway: IEEE; 2025. p. 5472-81.

[62]

Wu S, Hu C, Tian B, Huang Y, Yang S, Li S, et al. A 3D reconstruction platform for complex plants using OB—NeRF. Front Plant Sci 2025; 16: 1449626.

[63]

Cui W, Zhao C, Chen Y, Li M, Wang H . CL3R: 3D reconstruction and contrastive learning for enhanced robotic manipulation representations. 2025. arXiv:2502.11241.

[64]

Huang Y, Xu S, Chen H, Li G, Dong H, Yu J, et al. A review of visual perception technology for intelligent fruit harvesting robots. Front Plant Sci 2025; 16: 1646871.

[65]

Khan MN, Rahi A, Rajendran VP, Al Hasan M, Anwar S . Real time crop row detection using computer vision—application in agricultural robots. Front Artif Intell 2024; 7: 1435686.

[66]

Mac TT, Nguyen TD, Dang HK, Nguyen DT, Nguyen XT . Intelligent agricultural robotic detection system for greenhouse tomato leaf diseases using soft computing techniques. Sci Rep 2024; 14(1): 23887.

[67]

Wang R, Chen L, Huang Z, Wu S . A review on the high efficiency detection and precision positioning technology application of agricultural robots. Processes 2024; 12(9): 1833.

[68]

Tang C, Huang D, Dong W, Xu R, Zhang H . Foundationgrasp: generalizable task—oriented grasping with foundation models. IEEE Trans Autom Sci Eng 2025; 22: 12418-35.

[69]

Zhang C, Jiang D, Wan T, Rao Y, Jin X, Wang T, et al. Hallucination alleviation—based smart decision for early soybean cultivation in greenhouse and field scenarios. Comput Electron Agric 2025; 238: 110811.

[70]

Shaikh TA, Rasool T, Veningston K, Yaseen SM . The role of large language models in agriculture: harvesting the future with LLM intelligence. Prog Artif Intell 2025; 14(2): 117-64.

[71]

Rezayi S, Liu Z, Wu Z, Dhakal C, Ge B, Zhen C, et al. AgriBERT: knowledge infused agricultural language models for matching food and nutrition. In: Proceedings of the 31st International Joint Conference on Artificial Intelligence; 2022 Jul 23—29; Vienna, Austria. Washington: IJCAI; 2022. p. 5150—6.

[72]

Bo Y, Yu Z, Lan F, Yun C, Jian Z, Xiao X, et al. AgriGPT: a large language model ecosystem for agriculture. 2025. arXiv:2503.10189.

[73]

Mohan GB, Adhitya CJ, Mithilesh A . Enhancing agricultural advisory services with multilingual LLaMA and RAG. In: Proceedings of the 2025 International Conference on Next Generation Communication & Information Processing; 2025 Mar 21—23; Bangalore, India. Piscataway: IEEE; 2025. p. 550-6.

[74]

Fazlollahtabar H. Human—robot interaction using retrieval augmented generation and fine tuning with transformer neural networks in industry 5.0. Sci Rep 2025; 15(1): 29233.

[75]

Zhao B, Jin W, Del Ser J, Yang G . ChatAgri: exploring potentials of ChatGPT on cross linguistic agricultural text classification. Neurocomputing 2023; 557: 126708.

[76]

Atuhurra J. Enrich robots with updated knowledge in the wild via large language models. In: Proceedings of the ACM Conference 2024; 2024 Mar 13—17; held virtually. New York City: Association for Computing Machinery (ACM); 2024. p. 1-3.

[77]

Huang W, Wang C, Zhang R, Li Y, Fei Fei L . Composable 3d value maps for robotic manipulation with language models. 2023. arXiv:2306.14898.

[78]

Vemprala SH, Bonatti R, Bucker A, Kapoor A . ChatGPT for robotics: design principles and model abilities. IEEE Access 2024; 12: 55682-96.

[79]

Jiang Y, Gupta Z, Zhang Z, Wang G, Dou Y, Chen Y, et al. VIMA: general robot manipulation with multimodal prompts. In: Proceedings of the International Conference on Learning Representations 2023; 2023 Jul 23—29; Honolulu, HI, USA. London: PMLR; 2023. p. 1-48.

[80]

Castagna F, McBurney P, Parsons S . Explanation—question—response dialogue: an argumentative tool for explainable AI. Argument Comput 2025; 15(2): 133—50.

[81]

Ren AZ, Dixit A, Bodrova A, Singh S, Tu S, Brown N, et al. Robots that ask for help: uncertainty alignment for large language model planners. In: Agrawal P, Kroemer O, Burgard W, editors. Proceedings of the 7th Annual Conference on Robot Learning; 2023 Nov 6—9; Atlanta, GA, USA. London: PMLR; 2023.

[82]

Chi C, Xu Z, Feng S, Cousineau E, Du Y, Xu Z, et al. Diffusion policy: visuomotor policy learning via action diffusion. Int J Robot Res 2025; 44(10—11): 1684-704.

[83]

Park JH, Choi W, Hong S, Seo H, Ahn J, Ha C, et al. Hierarchical action chunking transformer: learning temporal multimodality from demonstrations with fast imitation behavior. In: Proceedings of the 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems; 2024 Oct 14—18; Abu Dhabi, UAE. Piscataway: IEEE; 2024. p. 12648—54.

[84]

Bar A, Zhou G, Tran D, Darrell T, LeCun Y . Navigation world models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; 2025 Jun 10—14; Nashville, TN, USA. Piscataway: IEEE; 2025. p. 15791—801.

[85]

Oquab M, Darcet T, Moutakanni T, Vo H, Szafraniec M, Piguet R, et al. DINOv2: learning robust visual features without supervision. 2024. arXiv:2304.07193v2.

[86]

Touvron H, Martin L, Stone K, Albert P, Almahairi A, Babaei Y, et al. Llama 2: open foundation and fine tuned chat models. 2023. arXiv:2307.09288.

[87]

Huang H, Liu F, Fu L, Wu T, Mukadam M, Malik J, et al. Early fusion helps vision language action models generalize better. In: Proceedings of the 1st Workshop on X Embodiment Robot Learning; 2024 Nov 6—9; Munich, Germany. Amherst: OpenReview; 2024.

[88]

Lu J, Clark C, Lee S, Zhang Z, Khosla S, Marten R, et al. Unified IO 2: scaling autoregressive multimodal models with vision, language, audio, and action. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; 2024 Jun 17—21; Seattle, WA, USA. Piscataway: IEEE; 2024. p. 26429—45.

[89]

Li Q, Liang Y, Wang Z, Luo L, Chen X, Liao M, et al. Cogact: a foundational vision—language—action model for synergizing cognition and action in robotic manipulation. 2024. arXiv:2410.02673.

[90]

Mirbod O, Choi D, Schueller JK . From simulation to field validation: a digital twin driven sim2real transfer approach for strawberry fruit detection and sizing. AgriEngineering 2025; 7(3): 81.

[91]

Williams E, Polydoros A . Zero shot sim to real reinforcement learning for fruit harvesting. In: Proceedings of the 2025 IEEE 21st International Conference on Automation Science and Engineering; 2025 Aug 17—21; Los Angeles, CA, USA. Piscataway: IEEE; 2025. p. 1423—8.

[92]

Mazumder A, Sahed MF, Tasneem Z, Das P, Badal FR, Ali MF, et al. Towards next generation digital twin in robotics: trends, scopes, challenges. Heliyon 2023; 9(9): e13359.

[93]

Zhou H, Kang H, Wang X, Au W, Wang MY, Chen C . Branch interference sensing and handling by tactile enabled robotic apple harvesting. Agronomy 2023; 13(2): 503.

[94]

Park S, Kim H, Kim H, Choi J . Pruning with scaled policy constraints for light weight reinforcement learning. IEEE Access 2024; 12: 36055—65.

[95]

Bharadhwaj H, Vakil J, Sharma M, Gupta A, Tulsiani S, Kumar V . Roboagent: generalization and efficiency in robot manipulation via semantic augmentations and action chunking. In: Proceedings of the 2024 IEEE International Conference on Robotics and Automation; 2024 May 13—17; Yokohama, Japan. Piscataway: IEEE; 2024. p. 4788—95.

[96]

Zhang B, Zhang Y, Ji J, Lei Y, Dai J, Chen Y, et al. SafeVLA: towards safety alignment of vision language action model via constrained learning. Adv Neural Inf Process Syst 2026; 38: 153335-73.

[97]

Angelopoulos AN, Bates S . Conformal prediction: a gentle introduction. Found Trends Mach Learn 2023; 16(4): 494-591.

[98]

Lindemann L, Cleaveland M, Shim G, Pappas GJ . Safe planning in dynamic environments using conformal prediction. IEEE Robot Autom Lett 2023; 8(8): 4634—41.

[99]

Yang Y, Caluwaerts K, Iscen A . Data efficient reinforcement learning for legged robots. In: Agrawal P, Kroemer O, Burgard W, editors. Proceedings of the 3rd Conference on Robot Learning; 2020 Nov 16—18; held virtually. London: PMLR; 2020. p. 1-10.

[100]

Miki T, Lee J, Hwangbo J, Wellhausen L, Koltun V, Hutter M . Learning robust perceptive locomotion for quadrupedal robots in the wild. Sci Robot 2022; 7(62): eabk2822.

[101]

Ma C, Ying Y, Xie L . Visuo tactile sensor development and its application for non destructive measurement of peach firmness. Comput Electron Agric 2024; 218: 108709.

[102]

Wang Y, Ding P, Li L, Cui C, Ge Z, Tong X, et al. VLA Adapter: an effective paradigm for tiny scale vision—language—action model. In: Proceedings of the AAAI Conference on Artificial Intelligence; 2026 Jan 20—27; Singapore City, Singapore. Washington, DC: AAAI Press; 2026. p. 18638-46.

[103]

Lu Y, Young S . A survey of public datasets for computer vision tasks in precision agriculture. Comput Electron Agric 2020; 178: 105760.

[104]

Hu EJ, Shen Y, Wallis P, Allen Zhu Z, Li Y, Wang S, et al. LoRA: low rank adaptation of large language models. 2022. arXiv:2106.09685.

[105]

Budzianowski P, Maa W, Freed M, Mo J, Hsiao W, Lee A, et al. EdgeVLA: efficient vision—language—action models. 2025. arXiv:2504.02177.

[106]

Yu P, Teng F, Zhu W, Shen C, Chen Z, Song J . Cloud—edge—device collaborative computing in smart agriculture: architectures, applications, and future perspectives. Front Plant Sci 2025; 16: 1668545.

PDF (8445KB)

0

Accesses

0

Citation

Detail

Sections
Recommended

/

〈 〉