A Software–Hardware Co-optimization and Evaluation Framework for Modular End-to-End Autonomous Driving

Chengzhi Ji , Xingfeng Li , Zhaodong Lv , Hao Sun , Pan Liu , Hao Frank Yang , Ziyuan Pu

Engineering ›› : 202607031

PDF (2226KB)
Engineering ›› :202607031 DOI: 10.1016/j.eng.2026.07.031
research-article
A Software–Hardware Co-optimization and Evaluation Framework for Modular End-to-End Autonomous Driving
Author information +
History +
PDF (2226KB)

Abstract

Modular end-to-end (ME2E) autonomous driving paradigms have achieved considerable closed-loop driving system performance improvements through continuous advancements in computational design. However, practical system deployment performance depends not only on accuracy but also on inference latency and energy consumption. To our knowledge, most existing studies still emphasize accuracy while giving limited consideration to deployment-oriented metrics. In addition, existing optimization methods remain largely confined to one-sided advances on either the software or hardware side, making it difficult to balance these multi-dimensional metrics and ultimately limiting real-world performance. To address these limitations, this paper proposes a generic software–hardware co-optimization and closed-loop evaluation framework for ME2E autonomous driving inference. The framework integrates software-side model compression with hardware-side computation acceleration under a unified objective. Furthermore, a multi-dimensional evaluation metric is introduced to jointly assess the performance of autonomous driving systems from both algorithmic and deployment-oriented dimensions. Experiments involving optimization and quantitative evaluation on multiple ME2E autonomous driving stacks show that the proposed framework preserves baseline-level accuracy while reducing inference latency by more than 6× and per-frame energy consumption to approximately one-fifth of the baseline. Consequently, a 22.35% improvement in the multi-dimensional performance evaluation index for autonomous vehicles (MPEAV) is achieved. The experimental results demonstrate that the proposed framework provides effective optimization guidance from both the software and hardware perspectives.

Keywords

Modular end-to-end autonomous driving / Software–hardware co-optimization / multi-dimensional closed-loop evaluation / Real-time synchronous simulation

Cite this article

Download citation ▾
Chengzhi Ji, Xingfeng Li, Zhaodong Lv, Hao Sun, Pan Liu, Hao Frank Yang, Ziyuan Pu. A Software–Hardware Co-optimization and Evaluation Framework for Modular End-to-End Autonomous Driving. Engineering 202607031 DOI:10.1016/j.eng.2026.07.031

登录浏览全文

4963

注册一个新账户 忘记密码

References

[1]

Hu J, Wang Y, Cheng S, Xu J, Wang N, Fu B, et al. A survey of decision—making and planning methods for self—driving vehicles. Front Neurorobot 2025; 19:1451923.

[2]

Hu Z, Xu M, Cheng Q . Multimodal large—language model empowering next—generation autonomous driving systems. J Intell Connect Veh 2025; 8(2):9210059.

[3]

Chen L, Wu P, Chitta K, Jaeger B, Geiger A, Li H . End—to—end autonomous driving: challenges and frontiers. IEEE Trans Pattern Anal Mach Intell 2024; 46(12):10164-83.

[4]

Jiang S, Huang Z, Qian K, Luo Z, Zhu T, Zhong Y, et al. A survey on vision—language—action models for autonomous driving. 2025. arXiv:2506.24044.

[5]

Wang H, Shao W, Sun C, Yang K, Cao D, Li J . A survey on an emerging safety challenge for autonomous vehicles: safety of the intended functionality. Engineering 2024; 33:17-34.

[6]

Ghoneim O, Dobias P, Romain O . Survey of neural network optimization methods for sustainable ai: from data preprocessing to hardware acceleration. Mach Learn Appl 2025; 22:100762.

[7]

Seif HG, Hu X . Autonomous driving in the iCity—HD maps as a key challenge of the automotive industry. Engineering 2016;2(2):159-62.

[8]

Sun K, Wang M, Wang Z . RETA—AD: a reconfigurable and efficient transformer accelerator for autonomous driving. IEEE Trans Very Large Scale Integr (VLSI) Syst 2025; 33(7):1945-58.

[9]

Xia C, Zhao J, Sun Q, Wang Z, Wen Y, Yu T, et al. Optimizing deep learning inference via global analysis and tensor expressions. In: Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems; 2024 Apr 27—May 1; La Jolla, CA, USA. New York City: Association for Computing Machinery; 2024. p. 286-301.

[10]

Deng L, Li G, Han S, Shi L, Xie Y . Model compression and hardware acceleration for neural networks: a comprehensive survey. Proc IEEE 2020; 108(4):485-532.

[11]

Zhang H, Xing M, Wu Y, Zhao C . Compiler technologies in deep learning co—design: a survey. Intell Comput 2023; 2:0040.

[12]

Liu L, Wang P, Wu G, Jiang J, Yang HF . Toward optimal mixture of experts system for 3D object detection: a game of accuracy, efficiency and adaptivity. IEEE Trans Pattern Anal Mach Intell 2026; 48(1):914—31.

[13]

Zhang H, Liu L, Hui F, Zhang B, Zhang H, Zha Z . Clean: category knowledge—driven compression framework for efficient 3D object detection. IEEE Trans Pattern Anal Mach Intell 2025; 47(10):8740-55.

[14]

Hu Y, Yang J, Chen L, Li K, Sima C, Zhu X, et al. Planning—oriented autonomous driving. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; 2023 Jun 17—24; Vancouver, BC, Canada. Piscataway: Institute of Electrical and Electronics Engineers; 2023. p. 17853—62.

[15]

Jiang B, Chen S, Xu Q, Liao B, Chen J, Zhou H, et al. VAD: Vectorized scene representation for efficient autonomous driving, in: Proceedings of the IEEE/CVF International Conference on Computer Vision; 2023 Oct 1—6; Paris, France. Piscataway: Institute of Electrical and Electronics Engineers; 2023. p. 8306—16.

[16]

Jia X, Yang Z, Li Q, Zhang Z, Yan J . Bench2drive: Towards multi—ability benchmarking of closed—loop end—to—end autonomous driving. In: Proceedings of the Advances in Neural Information Processing Systems, Vol. 37; 2024 Dec 9—15; Vancouver, BC, Canada. Red Hook: Curran Associates, Inc.; 2024. p. 819-44.

[17]

Yang Y, Zhan J, Liu Y, Wang Q . Cross—city transfer learning: Applications and challenges for smart cities and sustainable transportation. Commun Trans Res 2025; 5:100206.

[18]

Hussain F, Li Y, Haque MM . Integrating machine learning and extreme value theory for estimating crash frequency—by—severity via ai—based video analytics. Commun Trans Res 2024; 4:100147.

[19]

Huang Z, Sheng Z, Ma C, Chen S . Human as ai mentor: Enhanced human—in—the—loop reinforcement learning for safe and efficient autonomous driving. Commun Trans Res 2024; 4:100127.

[20]

Gao Y, Levinson D . Lane changing and congestion are mutually reinforcing? Commun Trans Res 2023; 3:100101.

[21]

Zhong C, Wu P, Zhang Q, Ma Z . Online prediction of network—level public transport demand based on principle component analysis. Commun Trans Res 2023; 3:100093.

[22]

Wu J, Sanchez—Diaz I, Yang Y, Qu X . Why is your paper rejected? lessons learned from over 5000 rejected transportation papers. Commun Trans Res 2024; 4:100129.

[23]

Tampuu A, Matiisen T, Semikin M, Fishman D, Muhammad N . A survey of end—to—end driving: architectures and training methods. IEEE Trans Neural Netw Learn Syst 2022; 33(4):1364-84.

[24]

Chitta K, Prakash A, Jaeger B, Yu Z, Renz K, Geiger A . Transfuser: Imitation with transformer—based sensor fusion for autonomous driving. IEEE Trans Pattern Anal Mach Intell 2023; 45(11):12878-95.

[25]

Prakash A, Chitta K, Geiger A . Multi—modal fusion transformer for end—to—end autonomous driving. In: Proceedings of the 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2021 Jun 19—25; Nashville, TN, USA. Piscataway: Institute of Electrical and Electronics Engineers; 2021. p. 7073-83.

[26]

Chen D, Krähenbühl P . Learning from all vehicles. In: Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2022 Jun 18—24; New Orleans, LA, USA. Piscataway: Institute of Electrical and Electronics Engineers; 2022. p. 17201—10.

[27]

Wu P, Jia X, Chen L, Yan J, Li H, Qiao Y . Trajectory—guided control prediction for end—to—end autonomous driving: A simple yet strong baseline. In: Proceedings of the Advances in Neural Information Processing Systems, Vol. 35; 2022 Nov 28—Dec 9; New Orleans, LA, USA. Red Hook: Curran Associates, Inc.; 2022. p. 6119—32.

[28]

Jia X, Wu P, Chen L, Xie J, He C, Yan J, et al. Think twice before driving: Towards scalable decoders for end—to—end autonomous driving. In: Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2023 Jun 17—24; Vancouver, BC, Canada. Piscataway: Institute of Electrical and Electronics Engineers; 2023. p. 21983—94.

[29]

Sadat A, Casas S, Ren M, Wu X, Dhawan P, Urtasun R . Perceive, predict, and plan: Safe motion planning through interpretable semantic representations. In: Proceedings of the Computer Vision — ECCV 2020; 2020 Aug 23—28; Glasgow, UK. Cham: Springer International Publishing; 2020. p. 414—30.

[30]

Hu S, Chen L, Wu P, Li H, Yan J, Tao D . ST—P3: End—to—end vision—based autonomous driving via spatial—temporal feature learning. In: Proceedings of the European Conference on Computer Vision; 2022 Oct 23—27; Tel Aviv, Israel. Cham: Springer Nature Switzerland; 2022. p. 533-49.

[31]

Gholami A, Kim S, Dong Z, Yao Z, Mahoney MW, Keutzer K . A survey of quantization methods for efficient neural network inference. In: Low—Power Computer Vision: Improve the Efficiency of Artificial Intelligence. Boca Raton: Chapman and Hall/CRC; 2022. p. 291-326.

[32]

Shao Z, Wang F, Sun T, Yu C, Fang Y, Jin G, et al. HUTFormer: hierarchical U—Net transformer for long—term traffic forecasting. Commun Trans Res 2025; 5:100218.

[33]

Shuvo MMH, Islam SK, Cheng J, Morshed BI . Efficient acceleration of deep learning inference on resource—constrained edge devices: a review. Proc IEEE 2023; 111(1):42-91.

[34]

Chen S, Jiang B, Gao H, Liao B, Xu Q, Zhang Q, et al. VADv2: end—to—end vectorized autonomous driving via probabilistic planning. 2024. arXiv:2402.13243.

[35]

Sun W, Lin X, Shi Y, Zhang C, Wu H, Zheng S . SparseDrive: end—to—end autonomous driving via sparse scene representation. In: Proceedings of the 2025 IEEE International Conference on Robotics and Automation (ICRA); 2025 May 19—23; Seattle, WA, USA. Piscataway: Institute of Electrical and Electronics Engineers; 2025. p. 8795-801.

[36]

Zheng W, Song R, Guo X, Zhang C, Chen L . GenAD: generative end—to—end autonomous driving. In: Proceedings of the European Conference on Computer Vision; 2024 Sep 29—Oct 4; Milan, Italy. Cham: Springer Nature Switzerland; 2024. p. 87-104.

[37]

Jia X, You J, Zhang Z, Yan J . DriveTransformer: Unified transformer for scalable end—to—end autonomous driving. In: Proceedings of the 2025 International Conference on Learning Representations (ICLR); 2025 May 5—9; Vienna, Austria. Red Hook: Curran Associates, Inc.; 2025. p. 67227-43.

[38]

Weng X, Ivanovic B, Wang Y, Wang Y, Pavone M . Para—Drive: Parallelized architecture for real—time autonomous driving. In: Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2024 Jun 17—21; Seattle, WA, USA. Piscataway: Institute of Electrical and Electronics Engineers; 2024. p. 15449—58.

[39]

Li Z, Yu Z, Lan S, Li J, Kautz J, Lu T, et al. Is ego status all you need for open—loop end—to—end autonomous driving? In: Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2024 Jun 17—21; Seattle, WA, USA. Piscataway: Institute of Electrical and Electronics Engineers;2024. p. 14864—73.

[40]

Zhai JT, Feng Z, Du J, Mao Y, Liu JJ, Tan Z, et al. Rethinking the open—loop evaluation of end—to—end autonomous driving in nuscenes. 2023. arXiv:2305.10430.

[41]

Dauner D, Hallgarten M, Geiger A, Chitta K . Parting with misconceptions about learning—based vehicle motion planning. In: Proceedings of the 2023 Conference on Robot Learning (CoRL)/Proceedings of Machine Learning Research (PMLR); 2023 Nov 6—9; Atlanta, GA, USA. Cambridge: MIT Press; 2023. p. 1268—81.

[42]

Gulino C, Fu J, Luo W, Tucker G, Bronstein E, Lu Y, et al. Waymax: An accelerated, data—driven simulator for large—scale autonomous driving research. In: Advances in Neural Information Processing Systems (NeurIPS), Vol. 36. New Orleans, LA, USA: Curran Associates, Inc.;2023. p. 7730—42.

[43]

Karnchanachari N, Geromichalos D, Tan KS, Li N, Eriksen C, Yaghoubi S, et al. Towards learning—based planning: The nuplan benchmark for real—world autonomous driving. In: Proceedings of the 2024 IEEE International Conference on Robotics and Automation (ICRA); 2024 May 13—17; Yokohama, Japan. Piscataway: Institute of Electrical and Electronics Engineers; 2024. p. 629—36.

[44]

Dosovitskiy A, Ros G, Codevilla F, Lopez A, Koltun V . CARLA: An open urban driving simulator. In: Proceedings of the 2017 Conference on Robot Learning (CoRL); 2017 Nov 13—15; Mountain View, CA, USA. Proceedings of Machine Learning Research (PMLR), Vol. 78. Cambridge: MIT Press; 2017. p. 1-16.

[45]

Renz K, Chitta K, Mercea OB, Koepke AS, Akata Z, Geiger A, et al. Explainable planning transformers via object—level representations. In: Proceedings of the 6th Conference on Robot Learning (CoRL); 2023 Nov 6—9; Atlanta, GA, USA. Proceedings of Machine Learning Research (PMLR), Vol. 205. Cambridge: MIT Press; 2023. p. 459-70.

[46]

Bender G, Kindermans PJ, Zoph B, Vasudevan V, Le Q . Understanding and simplifying one—shot architecture search. In: Proceedings of the 35th International Conference on Machine Learning (ICML); 2018 Jul 10—15; Stockholm, Sweden. Proceedings of Machine Learning Research (PMLR), Vol. 80. Cambridge: MIT Press; 2018. p. 550—9.

[47]

Zhao SZ, Zhang H, Li Z, Peng J, Chui A, Zhou Z, et al. Quantv2x: a fully quantized multi—agent system for cooperative perception. 2025. arXiv:2509.03704.

[48]

NVIDIA Corporation. TensorRT developer guide [Internet]. Santa Clara:NVIDIA Corporation ; undated [cited 2026 Jun 16]. Available from:https://docs.nvidia.com/deeplearning/tensorrt/developer—guide/index.html.

[49]

Almutairi AA, LeBlanc DJ, Kusari A . Car—stage: automated framework for large—scale high—dimensional simulated time—series data generation based on user—defined criteria. 2025. arXiv:2503.03100.

[50]

H. Zamani, L. Bhuyan, J. Chen, Z. Chen, GreenMD: energy—efficient matrix decomposition on heterogeneous multi—GPU systems. ACM Trans. Parallel Comput 2023; 10(2):1-25.

[51]

Tschand A, Rajan ATR, Idgunji S, Ghosh A, Holleman J, Kiraly C, et al. MLPerf power: benchmarking the energy efficiency of machine learning systems from watts to megawatts for sustainable AI. In: Proceedings of the 2025 IEEE International Symposium on High—Performance Computer Architecture (HPCA); 2025 Feb 22—26; San Diego, CA, USA. Piscataway: Institute of Electrical and Electronics Engineers;2025. p. 1201-16.

PDF (2226KB)

95

Accesses

0

Citation

Detail

Sections
Recommended

/