###
计算机系统应用英文版:2025,34(12):16-25
本文二维码信息
码上扫一扫!
基于RISC-V异构平台的大语言模型推理加速
(1.南京信息工程大学 软件学院, 南京 210044;2.中国科学院大学南京学院, 南京 211135;3.中国科学院 软件研究所, 北京 100190)
Inference Acceleration for Large Language Models on RISC-V Heterogeneous Platform
(1.School of Software, Nanjing University of Information Science & Technology, Nanjing 210044, China;2.University of Chinese Academy of Sciences, Nanjing, Nanjing 211135, China;3.Institute of Software, Chinese Academy of Sciences, Beijing 100190, China)
摘要
图/表
参考文献
相似文献
本文已被:浏览 579次   下载 1112
Received:April 30, 2025    Revised:June 24, 2025
中文摘要: 随着大语言模型在各类生成任务中的广泛应用, 其高计算负载对底层硬件平台提出了更高的性能要求. RISC-V 作为一种新兴的开源指令集架构, 凭借其良好的可定制性和扩展性, 展现出巨大的发展潜力. 然而在部署主流大模型方面, RISC-V平台仍面临生态不完善、算力受限等诸多挑战. 本文提出一种基于RISC-V平台的大语言模型推理加速方法, 通过构建寒武纪MLU370加速卡的异构运行环境, 成功完成了设备驱动移植、基础库编译与PyTorch框架适配. 在此基础上, 进一步设计了一种轻量级多线程优化策略, 提升注意力机制等核心算子在多核体系结构下的执行效率. 实验结果表明, 在SG2042+MLU370-S4平台上部署多个主流大模型时, 该方法在不依赖其他优化策略下, 实现最高达52.3倍的端到端推理加速, 验证了其在RISC-V异构平台上的可行性与通用性.
Abstract:With the widespread deployment of large language models (LLMs) across various generative tasks, their high computational demands impose stringent performance requirements on the underlying hardware. RISC-V, an emerging open-source instruction-set architecture, shows great potential owing to its excellent customizability and extensibility. Nevertheless, when deploying mainstream LLMs, the RISC-V ecosystem still faces challenges such as an incomplete software stack and limited compute capability. This study proposes an inference acceleration method for LLMs on RISC-V heterogeneous platforms. By establishing a heterogeneous runtime environment that integrates the Cambricon MLU370 accelerator, the device driver is ported, essential libraries are compiled, and the PyTorch framework is adapted. Building on this foundation, a lightweight multi-threading optimization scheme is further designed to improve the efficiency of core operators—especially the attention mechanism—on multi-core architectures. Experimental results on the SG2042+ MLU370-S4 platform show that, without relying on any additional optimizations, the proposed method achieves up to 52.3 times end-to-end inference speedup for several mainstream LLMs, thus demonstrating both the feasibility and broad applicability of the approach on RISC-V heterogeneous systems.
文章编号:     中图分类号:    文献标志码:
基金项目:国家重点研发计划(2023YFB4503902)
引用文本:
沈郑东,刘雨冬,于佳耕,田青.基于RISC-V异构平台的大语言模型推理加速.计算机系统应用,2025,34(12):16-25
SHEN Zheng-Dong,LIU Yu-Dong,YU Jia-Geng,TIAN Qing.Inference Acceleration for Large Language Models on RISC-V Heterogeneous Platform.COMPUTER SYSTEMS APPLICATIONS,2025,34(12):16-25