Current Status, Challenges, and Prospects for Machine Vision Empowered by Edge Intelligence

Bingrui Zhao , Yaonan Wang , Hui Zhang , Hai Wang , Kaiwen Tang , Xiangdong Liao , Yurong Chen , Xuesan Su , Ating Yin

Engineering ›› : 202607017

PDF (9982KB)
Engineering ›› :202607017 DOI: 10.1016/j.eng.2026.07.017
research-article
Current Status, Challenges, and Prospects for Machine Vision Empowered by Edge Intelligence
Author information +
History +
PDF (9982KB)

Abstract

As a core sensing and decision-making technology in industrial and automation domains, machine vision is required to maintain high-precision image-processing performance while simultaneously meeting multiple engineering constraints, such as real-time operation, miniaturization, and energy efficiency. These demands have driven the deep integration of machine vision and edge intelligence, giving rise to a new research frontier—real-time intelligent machine vision systems. This article systematically reviews the progress of this interdisciplinary field from three perspectives: algorithm evolution, hardware–software co-design, and engineering applications. First, it summarizes the development of intelligent vision algorithms represented by convolutional neural networks and vision transformers, emphasizing the role of model compression techniques—including pruning, quantization, and lightweight architecture design—in enabling edge deployment. Second, it explores edge computing platforms suitable for intelligent machine vision tasks and discusses typical strategies for hardware–software co-optimization in deploying artificial intelligence algorithms. Furthermore, through analyses of representative applications such as intelligent connected vehicles, urban surveillance systems, and robotics, the paper highlights the advantages and potential of edge vision systems in latency-sensitive scenarios. Finally, it identifies current challenges, including system integration, computational efficiency, energy efficiency, and intelligence level, and envisions emerging directions such as in-sensor computing, in-memory computing, neuromorphic computing, and on-device deployment of large models at the edge. Overall, this review emphasizes that the deep synergy among algorithms, hardware, and applications will be the key pathway toward realizing autonomous, adaptive, and perceptually intelligent machine vision empowered by edge intelligence.

Keywords

Edge computing / Machine vision / Artificial intelligence / Real-time systems / Embedded systems

Cite this article

Download citation ▾
Bingrui Zhao, Yaonan Wang, Hui Zhang, Hai Wang, Kaiwen Tang, Xiangdong Liao, Yurong Chen, Xuesan Su, Ating Yin. Current Status, Challenges, and Prospects for Machine Vision Empowered by Edge Intelligence. Engineering 202607017 DOI:10.1016/j.eng.2026.07.017

登录浏览全文

4963

注册一个新账户 忘记密码

References

[1]

Kalluri PR, Agnew W, Cheng M, Owens K, Soldaini L, Birhane A . Computer—vision research powers surveillance technology. Nature 2025; 643(8070):73-9.

[2]

Lavin A, Gilligan—Lee CM, Visnjic A, Ganju S, Newman D, Ganguly S, et al. Technology readiness levels for machine learning systems. Nat Commun 2022; 13(1):6039.

[3]

Pan W, Zheng J, Wang L, Luo Y . A future perspective on in—sensor computing. Engineering 2022; 14:19-21.

[4]

Wu M, Yu FR, Liu PX . Intelligence networking for autonomous driving in beyond 5G networks with multi—access edge computing. IEEE Trans Veh Technol 2022; 71(6):5853-66.

[5]

Tian Y, Wang J, Wang Y, Zhao C, Yao F, Wang X . Federated vehicular transformers and their federations: privacy—preserving computing and cooperation for autonomous driving. IEEE Trans Intell Veh 2022; 7(3):456-65.

[6]

Shi W, Zhang X, Wang Y, Zhang Q . Edge computing: state—of—the—art and future directions. J Comput Res Dev 2019; 56:69-89.

[7]

Zhou Z, Chen X, Li E, Zeng L, Luo K, Zhang J . Edge intelligence: paving the last mile of artificial intelligence with edge computing. Proc IEEE 2019; 107(8):1738-62.

[8]

Li E, Zeng L, Zhou Z, Chen X, Edge AI. On—demand accelerating deep neural network inference via edge computing. IEEE Trans Wirel Commun 2020; 19(1):447-57.

[9]

Sodhro AH, Pirbhulal S, De Albuquerque VHC . Artificial intelligence—driven mechanism for edge computing—based industrial applications. IEEE Trans Industr Inform 2019; 15(7):4235-43.

[10]

Gong T, Zhu L, Yu FR, Tang T . Edge intelligence in intelligent transportation systems: a survey. IEEE Trans Intell Transp Syst 2023; 24(9):8919—44.

[11]

El Jarroudi M, Kouadio L, Delfosse P, Bock CH, Mahlein AK, Fettweis X, et al. Leveraging edge artificial intelligence for sustainable agriculture. Nat Sustain 2024;7(7):846-54.

[12]

Sony Corporation . Sony releases two new global shutter sensors for vision systems [Internet]. Matawan:Vision Systems Design; 2024 Jun 24 [cited 2025 Nov 10]. Available from:https://www.vision—systems.com/cameras—accessories/image—sensors/article/55091825/sony—releases—two—new—global—shutter—sensors.

[13]

Sony Corporation . IMX238LQJ CMOS image sensor datasheet [Internet]. Brompton:DatasheetCafe; [cited 2025 Nov 10]. Available from:https://www.datasheetcafe.com/imx238—datasheet—cmos—image—sensor—sony/.

[14]

Liu S, Liu L, Tang J, Yu B, Wang Y, Shi W . Edge computing for autonomous driving: opportunities and challenges. Proc IEEE 2019; 107(8):1697-1716.

[15]

He K, Zhang X, Ren S, Sun J . Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; 2016 Jun 26—Jul 1; Las Vegas, NV, USA. Piscataway: IEEE; 2016. p. 770—8.

[16]

Dosovitskiy A . An image is worth 16x16 words: transformers for image recognition at scale. 2020. arXiv:201011929.

[17]

Mittal S, Vaishay S . A survey of techniques for optimizing deep learning on GPUs. J Systems Archit 2019; 99:101635.

[18]

Wang W, Li K, Ji B, Liu X, Yu J, Wu Q . A survey of AI inference technologies for on—device systems. IEEE Internet Things J 2025; 12(24):51927-51950.

[19]

Wu X, Zheng L, Liu C, Gao T . Hardware platform stripe noise estimation and random noise removal framework. IEEE Trans Geosci Remote Sens 2024; 62:1-13.

[20]

Guan J, Lai R, Lu Y, Li Y, Li H, Feng L, et al. Memory—efficient deformable convolution based joint denoising and demosaicing for UHD images. IEEE Trans Circuits Syst Video Technol 2022; 32(11):7346—58.

[21]

Yu C, Hou LZ . Realization of a real—time image denoising system for dashboard camera applications. IEEE Trans Consum Electron 2022; 68(2):181-90.

[22]

Ye J, Fu C, Cao Z, An S, Zheng G, Li B . Tracker meets night: a transformer enhancer for UAV tracking. IEEE Robot Autom Lett 2022; 7(2):3866—73.

[23]

Oh M, Kim CK, Kim B, Kang Y, Kim HG . Real—time terrain correction of satellite imagery—based solar irradiance maps using precomputed data and memory optimization. Remote Sens 2023; 15(16):3965.

[24]

Meng P, Zhang H, Wen H, Min D, Zhang Z, Zhang G . Design of real—time image processing system of medical high—definition electronic endoscope based on FPGA. In: Proceedings of the 2022 28th International Conference on Mechatronics and Machine Vision in Practice; 2022 Nov 17—19; Nanjing, China. Piscataway: IEEE; 2022. p. 1-6.

[25]

Liu Y, Li J, Huang K, Li X, Qi X, Chang L, et al. MobileSP: an FPGA—based real—time keypoint extraction hardware accelerator for mobile VSLAM. IEEE Trans Circuits Syst I Regul Pap 2022; 69(12):4919—29.

[26]

Vaidya B, Paunwala C . Lightweight hardware architecture for object detection in driver assistance systems. Int J Pattern Recognit Artif Intell 2022; 36(7):2250027.

[27]

Upadhyay BB, Sarawadekar K . VLSI design of saturation—based image dehazing algorithm. IEEE Trans Very Large Scale Integr VLSI Syst 2023; 31(7):959-68.

[28]

Martinez—Alpiste I, Golcarenarenji G, Wang Q, Alcaraz—Calero JM . Smartphone—based real—time object recognition architecture for portable and constrained systems. J Real—Time Image Process 2022; 19(1):103-15.

[29]

Gupta K, Singh A, Yeduri SR, Srinivas MB, Cenkeramaddi LR . Hand gestures recognition using edge computing system based on vision transformer and lightweight CNN. J Ambient Intell Humaniz Comput 2023; 14(3):2601—15.

[30]

Li J, Yan D, He F, Dong Z, Jiang M . A mixed—precision transformer accelerator with vector tiling systolic array for license plate recognition in unconstrained scenarios. IEEE Trans Intell Transp Syst 2024; 25(12):20280-94.

[31]

Alif MAR, Hussain M . Lightweight convolutional network with integrated attention mechanism for missing bolt detection in railways. Metrology 2024; 4(2):254-78.

[32]

Wickramasinghe S, Parikh D, Zhang B, Kannan R, Prasanna V, Busart C . VTR: an optimized vision transformer for SAR ATR acceleration on FPGA. In: Proceedings of the SPIE Defense + Commercial Sensing: Image Sensing Technologies: Materials, Devices, Systems, and Applications XI; 2024 Apr 21—26; Long Beach, CA, USA. Bellingham: SPIE; 2024. p. 138-53.

[33]

Gómez BF, Yi L, Ramalingam B, Rayguru MM, Hayat AA, Thejus P, et al. Deep learning based litter identification and adaptive cleaning using self—reconfigurable pavement sweeping robot. In: Proceedings of the 2022 IEEE 18th International Conference on Automation Science and Engineering; 2022 Aug 20—24; Mexico City, Mexico. Piscataway: IEEE; 2022. p. 2301—6.

[34]

De Prado M, Rusci M, Capotondi A, Donze R, Benini L, Pazos N . Robustifying the deployment of tinyml models for autonomous mini—vehicles. Sensors 2021; 21(4):1339.

[35]

La Salvia M, Torti E, Marenzi E, Danese G, Leporati F . Edge and cloud computing approaches in the early diagnosis of skin cancer with attention—based vision transformer through hyperspectral imaging. J Supercomput 2024; 80(11):16368-92.

[36]

He Z, Fan X, Peng Y, Shen Z, Jiao J, Liu M . Empointmovseg: sparse tensor—based moving—object segmentation in 3D lidar point clouds for autonomous driving—embedded system. IEEE Trans Comput Aided Des Integr Circuits Syst 2023; 42(1):41-53.

[37]

Li E, Zhou Z, Chen X . Edge intelligence: on—demand deep learning model co—inference with device—edge synergy. In: Proceedings of the 2018 Workshop on Mobile Edge Communications; 2018 Aug 20; Budapest, Hungary. New York City: Association for Computing Machinery; 2018. p. 31-6.

[38]

Han K, Wang Y, Chen H, Chen X, Guo J, Liu Z, et al. A survey on vision transformer. IEEE Trans Pattern Anal Mach Intell 2023; 45(1):87-110.

[39]

Meribout M, Baobaid A, Khaoua MO, Tiwari VK, Pena JP . State of art IoT and Edge embedded systems for real—time machine vision applications. IEEE Access 2022; 10:58287—301.

[40]

LeCun Y, Bengio Y, Hinton G . Deep learning. Nature 2015; 521(7553):436—44.

[41]

Jiao L, Wang D, Bai Y, Chen P, Liu F . Deep learning in visual tracking: a review. IEEE Trans Neural Netw Learn Syst 2023; 34(9):5497-516.

[42]

Li Z, Liu F, Yang W, Peng S, Zhou J . A survey of convolutional neural networks: analysis, applications, and prospects. IEEE Trans Neural Netw Learn Syst 2022; 33(12):6999-7019.

[43]

Brauwers G, Frasincar F . A general survey on attention mechanisms in deep learning. IEEE Trans Knowl Data Eng 2023; 35(4):3279-98.

[44]

Hu J, Shen L, Sun G . Squeeze—and—excitation networks. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; 2018 Jun 18—22; Salt Lake City, UT, USA. Los Alamitos: IEEE Computer Society; 2018. p. 7132—41.

[45]

Woo S, Park J, Lee JY, Kweon IS . CBAM: convolutional block attention module. In: Proceedings of the European Conference on Computer Vision; 2018 Sep 8—14; Munich, Germany. Cham: Springer; 2018. p. 3-19.

[46]

Liu Z, Lin Y, Cao Y, Hu H, Wei Y, Zhang Z, et al. Swin Transformer: hierarchical vision transformer using shifted windows. In: Proceedings of the IEEE/CVF International Conference on Computer Vision; 2021 Oct 11—17; Montreal, QC, Canada. Los Alamitos: IEEE Computer Society; 2021. p. 10012-22.

[47]

Wang W, Xie E, Li X, Fan DP, Song K, Liang D, et al. Pyramid vision transformer: a versatile backbone for dense prediction without convolutions. In: Proceedings of the IEEE/CVF International Conference on Computer Vision; 2021 Oct 11—17; Montreal, QC, Canada. Los Alamitos: IEEE Computer Society; 2021. p. 568-78.

[48]

Gulati A, Qin J, Chiu CC, Parmar N, Zhang Y, Yu J, et al. Conformer: convolution—augmented transformer for speech recognition. In: Proceedings of Interspeech 2020; 2020 Oct 25—29; Shanghai, China. Baixas: ISCA; 2020. p. 5036—40.

[49]

Woo S, Debnath S, Hu R, Chen X, Liu Z, Kweon IS, et al. Convnext v2: co—designing and scaling convnets with masked autoencoders. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; 2023 Jun 18—22; Vancouver, BC, Canada. Los Alamitos: IEEE Computer Society; 2023. p. 16133—42.

[50]

Deng L, Li G, Han S, Shi L, Xie Y . Model compression and hardware acceleration for neural networks: a comprehensive survey. Proc IEEE 2020; 108(4):485-532.

[51]

Han S, Pool J, Tran J, Dally W . Learning both weights and connections for efficient neural network. In: Proceedings of the Conference on Neural Information Processing Systems; 2015 Dec 7—12; Montreal, QC, Canada. San Diego: Neural Information Processing Systems Foundation; 2015. p. 1135—43.

[52]

Guo Y, Yao A, Chen Y . Dynamic network surgery for efficient DNNs. In: Proceedings of the Conference on Neural Information Processing Systems; 2016 Dec 5—10; Barcelona, Spain. San Diego: Neural Information Processing Systems Foundation; 2016. p. 1379-87.

[53]

Fujii T, Sato S, Nakahara H . A threshold neuron pruning for a binarized deep neural network on an FPGA. IEICE Trans Inf Syst 2018; E101(D):376-86.

[54]

Yu R, Li A, Chen CF, Lai JH, Morariu VI, Han X, et al. NISP: pruning networks using neuron importance score propagation. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; 2018 Jun 18—23; Salt Lake City, UT, USA. Los Alamitos: IEEE Computer Society; 2018. p. 9194-203.

[55]

Lu L, Jin Y, Bi H, Luo Z, Li P, Wang T, et al. Sanger: a co—design framework for enabling sparse attention using reconfigurable architecture. In: Proceedings of the 54th Annual IEEE/ACM International Symposium on Microarchitecture; 2021 Oct 18—22; online. New York City: Association for Computing Machinery; 2021. p. 977—91.

[56]

Libano F, Wilson B, Wirthlin M, Rech P, Brunhaver J . Understanding the impact of quantization, accuracy, and radiation on the reliability of convolutional neural networks on FPGAs. IEEE Trans Nucl Sci 2020; 67(7):1478—84.

[57]

NVIDIA. Achieving FP32 accuracy for INT8 inference using quantization aware training with NVIDIA TensorRT [Internet]. Santa Clara:NVIDIA Tech Blog; 2021 Aug 10 [cited 2025 Nov 10]. Available from:https://developer.nvidia.com/blog/achieving—fp32—accuracy—for—int8—inference—using—quantization—aware—training—with—tensorrt/.

[58]

Intel . Easily optimize deep learning with 8—bit quantization [Internet]. Santa Clara:Intel Communities Blog; 2022 Mar 9 [cited 2025 Nov 10]. Available from:https://community.intel.com/t5/Blogs/Tech—Innovation/Artificial—Intelligence—AI/Easily—Optimize—Deep—Learning—with—8—Bit—Quantization/post/1366571.

[59]

Zhao B, Wang Y, Zhang H, Zhang J, Chen Y, Yang Y . 4—bit CNN quantization method with compact LUT—based multiplier implementation on FPGA. IEEE Trans Instrum Meas 2023; 72:2008110.

[60]

Sun X, Wang N, Chen CY, Ni J, Agrawal A, Cui X, et al. Ultra—low precision 4—bit training of deep neural networks. Adv Neural Inf Process Syst 2020; 33:1796-807.

[61]

Yuan C, Agaian SS . A comprehensive review of binary neural network. Artif Intell Rev 2023; 56(11):12949-3013.

[62]

Rastegari M, Ordonez V, Redmon J, Farhadi A . XNOR—Net: ImageNet classification using binary convolutional neural networks. In: Proceedings of the 14th European Conference on Computer Vision; 2016 Oct 11—14; Amsterdam, The Netherlands. Cham: Springer International Publishing; 2016. p. 525—42.

[63]

Sandler M, Howard A, Zhu M, Zhmoginov A, Chen LC . Mobilenetv2: inverted residuals and linear bottlenecks. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; 2018 Jun 18—23; Salt Lake City, UT, USA. Los Alamitos: IEEE Computer Society; 2018. p. 4510—20.

[64]

Zhang X, Zhou X, Lin M, Sun J . ShuffleNet: an extremely efficient convolutional neural network for mobile devices. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; 2018 Jun 18—23; Salt Lake City, UT, USA. Los Alamitos: IEEE Computer Society; 2018. p. 6848—56.

[65]

Ma N, Zhang X, Zheng HT, Sun J . ShuffleNet v2: practical guidelines for efficient CNN architecture design. In: Proceedings of the European Conference on Computer Vision; 2018 Sep 8—14; Munich, Germany. Cham: Springer International Publishing; 2018. p. 116-31.

[66]

Ding X, Zhang X, Ma N, Han J, Ding G, Sun J . RepVGG: making VGG—style convnets great again. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; 2021 Jun 20—25; online. Los Alamitos: IEEE Computer Society; 2021. p. 13733—42.

[67]

Ding X, Zhang X, Han J, Ding G . Diverse branch block: building a convolution as an inception—like unit. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; 2021 Jun 20—25; Virtual. Los Alamitos: IEEE Computer Society; 2021. p. 10886-95.

[68]

Tan Z, Li X, Wu Y, Chu Q, Lu L, Yu N, et al. Boosting vanilla lightweight vision transformers via re—parameterization. In: Proceedings of the Twelfth International Conference on Learning Representations; 2024 May 7—11; Vienna, Austria. New York City: JMLR.org; 2024.

[69]

Li M, Liu Y, Liu X, Sun Q, You X, Yang H, et al. The deep learning compiler: a comprehensive survey. IEEE Trans Parallel Distrib Syst 2021; 32(3):708—27.

[70]

Meng J, Zhuang C, Chen P, Wahib M, Schmidt B, Wang X, et al. Automatic generation of high—performance convolution kernels on ARM CPUs for deep learning. IEEE Trans Parallel Distrib Syst 2022; 33(11):2885-99.

[71]

Li X, Liang Y, Yan S, Jia L, Li Y . A coordinated tiling and batching framework for efficient GEMM on GPUs. In: Proceedings of the 24th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming; 2019 Feb 16—20; Washington, DC, USA. New York City: Association for Computing Machinery; 2019. p. 229—41.

[72]

Jiang ZH, Fei Y, Kaeli D . Exploiting bank conflict—based side—channel timing leakage of GPUs. ACM Trans Archit Code Optim 2019; 16(4):1-24.

[73]

Chen YH, Krishna T, Emer JS, Sze V . Eyeriss: an energy—efficient reconfigurable accelerator for deep convolutional neural networks. IEEE J Solid—State Circuits 2017; 52(1):127-38.

[74]

Mittal S. A survey of FPGA—based accelerators for convolutional neural networks. Neural Comput Appl 2020; 32(4):1109—39.

[75]

Lyu Y, Bai L, Huang X . ChipNet: real—time LiDAR processing for drivable region segmentation on an FPGA. IEEE Trans Circuits Syst I Regul Pap 2019; 66(5):1769-79.

[76]

Lin YJ, Chang TS . Data and hardware efficient design for convolutional neural network. IEEE Trans Circuits Syst I Regul Pap 2018; 65(5):1642—51.

[77]

Jouppi NP, Young C, Patil N, Patterson D, Agrawal G, Bajwa R, et al. In—datacenter performance analysis of a tensor processing unit. In: Proceedings of the 44th ACM/IEEE International Symposium on Computer Architecture; 2017 Jun 24—28; Toronto, ON, Canada. New York City: Association for Computing Machinery; 2017. p. 1-12.

[78]

Wei X, Yu CH, Zhang P, Chen Y, Wang Y, Hu H, et al. Automated systolic array architecture synthesis for high throughput CNN inference on FPGAs. In: Proceedings of the 54th Annual Design Automation Conference; 2017 Jun 18—22; Austin, TX, USA. New York City: Association for Computing Machinery; 2017. p. 1-6.

[79]

Li Y, Liu Z, Liu W, Jiang Y, Wang Y, Goh WL, et al. A 34—FPS 698—GOP/s/W binarized deep neural network—based natural scene text interpretation accelerator for mobile edge computing. IEEE Trans Ind Electron 2019; 66(9):7407—16.

[80]

Zhao B, Wang Y, Zhang H, Liu J, Zhang J . KAM—Net: kilobyte—scale ultralightweight attention—based network for glass defect detection with algorithm/hardware co—design. IEEE Trans Instrum Meas 2025; 74:5041310.

[81]

Samajdar A, Joseph JM, Zhu Y, Whatmough P, Mattina M, Krishna T . A systematic methodology for characterizing scalability of DNN accelerators using scale—sim. In: Proceedings of the 2020 IEEE International Symposium on Performance Analysis of Systems and Software; 2020 Aug 23—25; Boston, MA, USA. New York City: Association for Computing Machinery; 2020. p. 58-68.

[82]

Moolchandani D, Kumar A, Sarangi SR . Accelerating CNN inference on ASICs: a survey. J Systems Archit 2021; 113:101887.

[83]

Huang M, Luo J, Ding C, Wei Z, Huang S, Yu H . An integer—only and group—vector systolic accelerator for efficiently mapping vision transformer on edge. IEEE Trans Circuits Syst I Regul Pap 2023; 70(12):5289—301.

[84]

Wang Z, Wang G, He G . COSA plus: enhanced co—operative systolic arrays for attention mechanism in transformers. IEEE Trans Comput Aided Des Integr Circuits Syst 2024; 44(2):723—36.

[85]

Li Z, Sun M, Lu A, Ma H, Yuan G, Xie Y, et al. Auto—vit—acc: an FPGA—aware automatic acceleration framework for vision transformer with mixed—scheme quantization. In: Proceedings of the 2022 32nd International Conference on Field—Programmable Logic and Applications; 2022 Aug 29—Sep 02; Belfast, United Kingdom. New York City: Association for Computing Machinery; 2022. p. 109-16.

[86]

Wang T, Gong L, Wang C, Yang Y, Gao Y, Zhou X, et al. Via: a novel vision—transformer accelerator based on FPGA. IEEE Trans Comput Aided Des Integr Circuits Syst 2022; 41(11):4088-99.

[87]

You H, Sun Z, Shi H, Yu Z, Zhao Y, Zhang Y, et al. Vitcod: vision transformer acceleration via dedicated algorithm and accelerator co—design. In: Proceedings of the 2023 IEEE International Symposium on High—Performance Computer Architecture (HPCA’23); 2023 Feb 25 — Mar 01; Montreal, QC, Canada. New York City: Association for Computing Machinery;2023. p. 273-86.

[88]

Chen C, Li L, Aly MMS . ViTA: a highly efficient dataflow and architecture for vision transformers. In: Proceedings of the 2024 Design, Automation and Test in Europe Conference and Exhibition; 2024 Mar 25—Mar 27; Valencia, Spain. New York City: Association for Computing Machinery; 2024. p. 1-6.

[89]

Dong P, Zhuang J, Yang Z, Ji S, Li Y, Xu D, et al. EQ—ViT: Algorithm—hardware co—design for end—to—end acceleration of real—time vision transformer inference on Versal ACAP architecture. IEEE Trans Comput Aided Des Integr Circuits Syst 2024; 43(11):3949-60.

[90]

Sun X, Zhang Y, Wang Q, Zou X, Liu Y, Zeng Z, et al. DRViT: A dynamic redundancy—aware vision transformer accelerator via algorithm and architecture co—design on FPGA. J Parallel Distrib Comput 2025; 199:105042.

[91]

Guo Q, Wan J, Xu S, Li M, Wang Y . HG—PIPE: vision transformer acceleration with hybrid—grained pipeline. In: Proceedings of the 43rd IEEE/ACM International Conference on Computer—Aided Design; 2024 Oct 20—24; Newark, NJ, USA. New York City: Association for Computing Machinery (ACM); 2024. p. 1-9.

[92]

Zhang W, Zhang Y, Liu Y, Wu L, Hu X . REATA: an efficient vision transformer accelerator featuring a resource—optimized attention design on versal ACAP. ACM Trans Reconfig Technol Syst 2026; 19(1):1-32.

[93]

Xu C, Kan Y, Zhang R, Nakashima Y . An FPGA accelerator for vision transformer with quantization and LUT—based operations. IEICE Trans Inf Syst 2025; E109—D:41-8.

[94]

Xu R, Razavi S, Zheng R . Edge video analytics: a survey on applications, systems and enabling techniques. IEEE Commun Surv Tutor 2023; 25(4):2951-82.

[95]

Pereira D, Ghosh S, Dey S . Multi—stream scheduling of inference pipelines on edge devices—a DRL approach. ACM Trans Des Autom Electron Syst 2024; 29(6):1-36.

[96]

Zhang L, Lu Z, Song L, Xu J . CrossVision: real—time on—camera video analysis via common RoI load balancing. IEEE Trans Mobile Comput 2024; 23(5):5027-39.

[97]

Nan Y, Jiang S, Li M . Large—scale video analytics with cloud—edge collaborative continuous learning. ACM Trans Sens Netw 2024; 20(1):1-23.

[98]

Ghosh SK, Raha A, Raghunathan V, Raghunathan A . PArtNNer: platform—agnostic adaptive edge—cloud DNN partitioning for minimizing end—to—end latency. ACM Trans Embed Comput Syst 2024; 23(1):1-38.

[99]

Ji X, Gong F, Wang N, Du C, Yuan X . Task offloading with enhanced Deep Q—Networks for efficient industrial intelligent video analysis in edge—cloud collaboration. Adv Eng Inform 2024; 62:102599.

[100]

Intel. “Data is the new oil:” Intel announces new investment $250M in autonomous driving [Internet]. Santa Clara: Intel;2016 Jun 17 [cited 2025 Nov 11]. Available from:https://mobility21.cmu.edu/data—is—the—new—oil—intel—announces—new—investment—250m—in—autonomous—driving/.

[101]

Samsung Semiconductor . Autonomous driving and the modern data center—why high—performance memory and storage solutions are essential [Internet]. Seoul:Samsung Newsroom; [cited 2025 Nov 11]. Available from:https://semiconductor.samsung.com/news—events/tech—blog/autonomous—driving—and—the—modern—data—center/.

[102]

Betz J, Zheng H, Liniger A, Rosolia U, Karle P, Behl M, et al. Autonomous vehicles on the edge: a survey on autonomous vehicle racing. IEEE Open J Intell Transp Syst 2022; 3:458-88.

[103]

Cui C, Ma Y, Cao X, Ye W, Zhou Y, Liang K, et al. A survey on multimodal large language models for autonomous driving. IEEE Trans Intell Veh 2024; 9(5):958-79.

[104]

Liu L, Lu S, Zhong R, Wu B, Yao Y, Zhang Q, et al. Computing systems for autonomous driving: state of the art and challenges. IEEE Internet Things J 2021;8(8):6469—86.

[105]

Zhang P. BYD, Xpeng join Zeekr, Li Auto in adopting Nvidia’s next—gen Thor chip [Internet]. Singapore:CnEVPost; 2024 Mar 19 [cited 2025 Nov 11]. Available from:https://cnevpost.com/2024/03/19/byd—xpeng—adopt—nvidia—thor/.

[106]

TheFifthDriver. Machine learning driving assistance on FPGA [Internet]. Montreal:Hackster.io; [cited 2025 Nov 11]. Available from:https://www.hackster.io/javier—cristian/thefifthdriver—machine—learning—driving—assistance—on—fpga—98f295.

[107]

Kendall A, Hawke J, Janz D, Mazur P, Reda D, Allen JM, et al. Learning to drive in a Day. In: Proceedings of the 2019 International Conference on Robotics and Automation; 2019 May 20—24; Montreal, QC, Canada. New York City: Association for Computing Machinery; 2019. p. 8248-54.

[108]

Mobileye. Mobileye launches the first camera—only intelligent speed assist to meet new EU Standards [Internet]. Jerusalem: Mobileye; [cited2025 Nov 11]. Available from:https://www.mobileye.com/news/mobileye—launches—the—first—camera—only—intelligent—speed—assist—to—meet—new—eu—standards/.

[109]

Eyes on the road. Enabling real—time traffic camera analysis with BrainChip Akida [Internet]. San Jose: Edge Impulse; [cited2025 Nov 11]. Available from: https:https://www.edgeimpulse.com/blog/eyes—on—the—road—enabling—real—time—traffic—camera—analysis—with—brainchip—akida/.

[110]

Liu G, Shi H, Kiani A, Khreishah A, Lee J, Ansari N, et al. Smart traffic monitoring system using computer vision and edge computing. IEEE Trans Intell Transp Syst 2022; 23(8):12027—38.

[111]

Nikodem M, Słabicki M, Surmacz T, Mrówka P, Dołęga C . Multi—camera vehicle tracking using edge computing and low—power communication. Sensors 2020; 20(11):3334.

[112]

Wan S, Ding S, Chen C . Edge computing enabled video segmentation for real—time traffic monitoring in internet of vehicles. Pattern Recognit 2022; 121:108146.

[113]

AspenCore Media . Teardown: microbolometer—based intelligent thermal camera. Hong Kong: EE Times Asia; [cited 2025 Nov 11]. Available from:https://www.eetasia.com/teardown—microbolometer—based—intelligent—thermal—camera/.

[114]

Kong X, Wang K, Wang S, Wang X, Jiang X, Guo Y, et al. Real—time mask identification for COVID—19: an edge—computing—based deep learning framework. IEEE Internet Things J 2021; 8(21):15929-38.

[115]

Rahman MA, Hossain MS . An internet—of—medical—things—enabled edge computing framework for tackling COVID—19. IEEE Internet Things J 2021; 8(21):15847-854.

[116]

Yang J, Qian T, Zhang F, Khan SU . Real—time facial expression recognition based on edge computing. IEEE Access 2021; 9:76178-90.

[117]

Liu X, Yang J, Zou C, Chen Q, Yan X, Chen Y, et al. Collaborative edge computing with FPGA—based CNN accelerators for energy—efficient and time—aware face tracking system. IEEE Trans Comput Soc Syst 2022; 9(1):252-66.

[118]

Wu Y, Zhang L, Gu Z, Lu H, Wan S . Edge—AI—driven framework with efficient mobile network design for facial expression recognition. ACM Trans Embed Comput Syst 2023; 22(3):1-17.

[119]

Zhang J, Chen J, Wu W, Qiao H . A cerebellum—inspired prediction and correction model for motion control of a musculoskeletal robot. IEEE Trans Cogn Dev Syst 2023; 15(3):1209—23.

[120]

Qiao H, Chen J, Huang X . A survey of brain—inspired intelligent robots: integration of vision, decision, motion control, and musculoskeletal systems. IEEE Trans Cybern 2022; 52(10):11267—80.

[121]

Khomami AM, Najafi F . A survey on soft lower limb cable—driven wearable robots without rigid links and joints. Robot Auton Syst 2021; 144:103846.

[122]

AMD. Kria KR260 robotics starter kit [Internet]. Santa Clara: AMD; [cited2025 Nov 11]. Available from:https://www.amd.com/en/products/system—on—modules/kria/k26/kr260—robotics—starter—kit.html.

[123]

Hitachi, Ltd. Research & Development Group . Bringing together AI and FPGA technologies to overcome the shortcomings of robots [Internet]. Tokyo: Hitachi, Ltd.; [cited2025 Nov 11]. Available from:https://rd.hitachi.com/_ct/17711703.

[124]

Duan J, Yu S, Tan HL, Zhu H, Tan C . A survey of embodied AI: from simulators to research tasks. IEEE Trans Emerging Top Comput Intell 2022; 6(2):230-44.

[125]

Jin D, Zhang L . Embodied intelligence weaves a better future. Nat Mach Intell 2020; 2(11):663—4.

[126]

Bartolozzi C, Indiveri G, Donati E . Embodied neuromorphic intelligence. Nat Commun 2022; 13(1):1024.

[127]

Putra RVW, Shafique M . SpikeDyn: a framework for energy—efficient spiking neural networks with continual and unsupervised learning capabilities in dynamic environments. In: Proceedings of the 58th ACM/IEEE Design Automation Conference; 2021 Dec 5—9; San Francisco, CA, USA. New York City: Association for Computing Machinery; 2021. p. 1057-62.

[128]

Mazouz AE, Nguyen VT . Online continual streaming learning for embedded space applications. J Real—Time Image Process 2024; 21(3):68.

[129]

Tang Y, Zhang X, Zhou P, Hu J . EF—train: enable efficient on—device CNN training on FPGA through data reshaping for online adaptation or personalization. ACM Trans Des Autom Electron Syst 2022; 27(5):1-36.

[130]

Zitkovich B, Yu T, Xu S, Xu P, Xiao T, Xia F, et al. Rt—2: vision—language—action models transfer web knowledge to robotic control. In: Proceedings of the 6th Conference on Robot Learning; 2023 Nov 6—9; Atlanta, GA, USA. Cambridge: PMLR; 2023. p. 2165-83.

[131]

Kim MJ, Pertsch K, Karamcheti S, Xiao T, Balakrishna A, Nair S, et al. OpenVLA: an open—source vision—language—action model. In: Proceedings of the 8th Conference on Robot Learning; 2024 Nov 4—6; Osaka, Japan. Cambridge: PMLR;2025. p. 2679-713.

[132]

Shukor M, Aubakirova D, Capuano F, Kooijmans P, Palma S, Zouitine A, et al. SmolVLA: a vision—language—action model for affordable and efficient robotics. 2025. arXiv:2506.01844.

[133]

Jiang W, Clemons J, Sankaralingam K, Kozyrakis C . How fast can I run my VLA? Demystifying VLA inference performance with VLA—perf. 2026. arXiv:2602.18397.

[134]

Gubernatorov K, Voronov A, Voronov R, Pasynkov S, Perminov S, Guo Z, et al. AnywhereVLA: language—conditioned exploration and mobile manipulation. 2025. arXiv:2509.21006.

[135]

Li S, Song W, Fang L, Chen Y, Ghamisi P, Benediktsson JA . Deep learning for hyperspectral image classification: an overview. IEEE Trans Geosci Remote Sens 2019; 57(9):6690-709.

[136]

Misra NN, Dixit Y, Al—Mallahi A, Bhullar MS, Upadhyay R, Martynenko A . IoT, big data, and artificial intelligence in agriculture and food industry. IEEE Internet Things J 2022;9(9):6305—24.

[137]

Medus LD, Saban M, Francés—Víllora JV, Bataller—Mompeán M, Rosado—Muñoz A . Hyperspectral image classification using CNN: application to industrial food packaging. Food Control 2021; 125:107962.

[138]

Zhao B, Zhang H, Wang Y, Jiang Y, Chen Y, Yin A, et al. Unleashing the potential of on—chip ai—powered hyperspectral hardware computing—a tutorial. IEEE Trans Circuits Syst II Express Briefs 2024; 71(3):1733—7.

[139]

Rajesh R, Darak SJ, Jain A, Chandhok S, Sharma A . Hardware—software co—design of statistical and deep—learning frameworks for wideband sensing on zynq system on chip. IEEE Trans Very Large Scale Integr VLSI Syst 2023; 31(1):79-89.

[140]

Giuffrida G, Fanucci L, Meoni G, Batič M, Buckley L, Dunne A . The Φ—sat—1 mission: the first on—board deep neural network demonstrator for satellite earth observation. IEEE Trans Geosci Remote Sens 2022; 60(3):5517414.

[141]

Lo YC, Wu YC, Yang CHA . 44.3mW 62.4fps hyperspectral image processor for MAV remote sensing. In: Proceedings of the IEEE Symposium on VLSI Technology & Circuits; 2022 Jun 13—17; Honolulu, HI, USA. Piscataway: IEEE;2022. p. 74-5.

[142]

Zhang H, Wang H, Zhao B, Niu T, Cao Y, Wang Y . An integrated sensing—computing processor for spectral—based Chinese herbal medicine classification with algorithm—FPGA co—design. IEEE Trans Ind Electron 2025; 73(5):7995-8006.

[143]

Zhou F, Chai Y . Near—sensor and in—sensor computing. Nat Electron 2020; 3(11):664-71.

[144]

Ballard Z, Brown C, Madni AM, Ozcan A . Machine learning and computation—enabled intelligent sensor design. Nat Mach Intell 2021; 3(7):556-65.

[145]

Chen J, Wang W, Yan X . Two—dimensional material—based devices for in—sensor computing. npj Unconv Comput 2025; 2(1):19.

[146]

Fabre W, Haroun K, Lorrain V, Lepecq M, Sicard G . From near—sensor to in—sensor: a state—of—the—art review of embedded AI vision systems. Sensors 2024; 24(16):5446.

[147]

Wu W, Zhou T, Fang L . Parallel photonic chip for nanosecond end—to—end image processing, transmission, and reconstruction. Optica 2024; 11(6):831—7.

[148]

Ren Q, Zhu C, Ma S, Wang Z, Yan J, Wan T, et al. Optoelectronic devices for in—sensor computing. Adv Mater 2025; 37(23):2407476.

[149]

Bogaerts W, Pérez D, Capmany J, Miller DAB, Poon J, Englund D, et al. Programmable photonic circuits. Nature 2020; 586(7828):207—16.

[150]

Sebastian A, Le Gallo M, Khaddam—Aljameh R, Eleftheriou E . Memory devices and applications for in—memory computing. Nat Nanotechnol 2020; 15(7):529-44.

[151]

Yao P, Wu H, Gao B, Tang J, Zhang Q, Zhang W, et al. Fully hardware—implemented memristor convolutional neural network. Nature 2020; 577(7792):641-6.

[152]

Xia Q, Yang JJ . Memristive crossbar arrays for brain—inspired computing. Nat Mater 2019; 18(4):309—23.

[153]

Li Y, Tang J, Gao B, Yao J, Fan A, Yan B, et al. Monolithic three—dimensional integration of RRAM—based hybrid memory architecture for one—shot learning. Nat Commun 2023; 14(1):7140.

[154]

Wan W, Kubendran R, Schaefer C, Eryilmaz SB, Zhang W, Wu D, et al. A compute—in—memory chip based on resistive random—access memory. Nature 2022; 608(7923):504—12.

[155]

de la Rosa—Vidal R, Leñero—Bardallo JA, Gómez—Merchán R, Rodríguez—Vázquez Á . Event—driven vision sensor with in—pixel spatial contrast computation capabilities and on—chip AER sequencer. IEEE Trans Circuits Syst I Regul Pap 2025; 72(4):1234-48.

[156]

Gallego G, Delbrück T, Orchard G, Bartolozzi C, Taba B, Censi A, et al. Event—based vision: a survey. IEEE Trans Pattern Anal Mach Intell 2022; 44(1):154-80.

[157]

Brandli C, Berner R, Yang M, Liu SC, Delbruck TA . 240×180 130 db 3 μs latency global shutter spatiotemporal vision sensor. IEEE J Solid—State Circuits 2014; 49(10):2333—41.

[158]

Gehrig D, Gehrig M, Hidalgo—Carrió J, Scaramuzza D . Video to events: recycling video datasets for event cameras. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; 2020 Jun 14—19; Seattle, WA, USA. Los Alamitos: IEEE Computer Society; 2020. p. 3586-95.

[159]

Maass W. Networks of spiking neurons: the third generation of neural network models. Neural Netw 1997; 10(9):1659—71.

[160]

Bulzomi H, Memudu AS, Nakano Y, Martinet J . Real—time pedestrian detection at the edge on a fully asynchronous neuromorphic system. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; 2025 Jun 18—21; Melbourne, VIC, Australia. Los Alamitos: IEEE Computer Society;2025. p. 4958-67.

[161]

Ahmed SH, Finkbeiner J, Neftci E . Efficient event—based object detection: a hybrid neural network with spatial and temporal attention. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; 2025 Jun 18—21; Melbourne, VIC, Australia. Los Alamitos: IEEE Computer Society; 2025. p. 13970—9.

[162]

Setyawan N, Sun CC, Hsu MH, Kuo WK, Hsieh JW . MicroViT: a vision transformer with low complexity self attention for edge device. In: Proceedings of the 2025 IEEE International Symposium on Circuits and Systems; 2025 May 26—29; Montreal, QC, Canada. Piscataway: IEEE; 2025. p. 1-5.

[163]

Wu Y, Deng L, Li G, Zhu J, Shi L . Spatio—temporal backpropagation for training high—performance spiking neural networks. Front Neurosci 2018; 12:331.

[164]

Zhang M, Gu Z, Zheng N, Ma D, Pan G . Efficient spiking neural networks with logarithmic temporal coding. IEEE Access 2020; 8:98156—67.

[165]

Davies M, Wild A, Orchard G, Sandamirskaya Y, Guerra GAF, Joshi P, et al. Advancing neuromorphic computing with loihi: a survey of results and outlook. Proc IEEE 2021; 109(5):911—34.

[166]

Friha O, Amine Ferrag M, Kantarci B, Cakmak B, Ozgun A, Ghoualmi—Zine N . LLM—based edge intelligence: a comprehensive survey on architectures, applications, security and trustworthiness. IEEE Open J Commun Soc 2024; 5:5799-856.

[167]

Sun M, Liu Z, Bair A, Kolter JZ . A simple and effective pruning approach for large language models. 2024. arXiv:2306.11695.

[168]

Ma S, Wang H, Ma L, Wang L, Wang W, Huang S, et al. The era of 1—bit LLMs: all large language models are in 1.58 bits. 2024. arXiv:2402.17764.

[169]

Xiao G, Lin J, Seznec M, Wu H, Demouth J, Han S . SmoothQuant: accurate and efficient post—training quantization for large language models. In: Proceedings of the 40th International Conference on Machine Learning; 2023 Jul 23—29; Honolulu, HI, USA. Cambridge: PMLR; 2023. p. 38087—99.

[170]

Yi R, Guo L, Wei S, Zhou A, Wang S, Xu M . EdgeMoE: epowering sparse large language models on mobile devices. 2023. arXiv:2308.14352

[171]

Fan T, Kang Y, Ma G, Chen W, Wei W, Fan L, et al. FATE—LLM: a industrial grade federated learning framework for large language models. 2023. arXiv:2310.10049.

PDF (9982KB)

104

Accesses

0

Citation

Detail

Sections
Recommended

/