Parallel DSMC algorithm for unstructured grids on CUDA platform
-
摘要:
为提升稀薄气体流动模拟的计算效率,针对非结构网格在复杂几何建模中的适应性特点,构建了一种基于统一计算设备架构(CUDA)平台的非结构网格直接模拟蒙特卡洛(DSMC)并行算法。建立通用的DSMC求解框架,设计高效的网格搜索与粒子追踪机制,并在CUDA平台上对流场初始化、粒子推进、碰撞、排序及结果取样等核心模块进行了并行优化。通过合理划分线程与存储资源,解决了图形处理器(GPU)计算中数据一致性、线程发散与负载不均衡等问题。以超声速圆柱绕流为算例进行验证,结果表明:该算法在单卡GPU上整体加速比达到52.5;基于网格并行策略的各模块取得了24.7~53.7的加速比,基于分子并行策略的各模块取得了47~68倍的加速比。该算法显著提升了非结构网格DSMC模拟的并行性能,为高马赫数稀薄气体流动的高效数值模拟提供了一种可行方案。
-
关键词:
- 直接模拟蒙特卡洛(DSMC)方法 /
- 非结构网格 /
- 图形处理器(GPU)并行计算 /
- 统一计算设备架构(CUDA)平台 /
- 稀薄气体动力学
Abstract:To enhance the computational efficiency of rarefied gas flow simulations, a compute unified device architecture (CUDA) platform-based parallel direct simulation Monte Carlo (DSMC) algorithm for unstructured grids was developed considering the adaptability of unstructured meshes to complex geometries. A general DSMC framework was established, and efficient grid-search and particle-tracking mechanisms were designed. The core computational modules—including flow-field initialization, particle movement, collision, sorting, and sampling—were parallelized and optimized on the CUDA platform. By rationally allocating threads and memory resources, issues such as data consistency, thread divergence, and load imbalance in graphics processing unit (GPU) computing were effectively addressed. The algorithm was validated through simulation of supersonic flow over a circular cylinder. Results showed that the overall speedup on a single GPU reached 52.5; the speedups of individual modules based on the grid-parallel strategy ranged from 24.7 to 53.7, while those based on the molecule-parallel strategy ranged from 47 to 68. The proposed algorithm significantly improved the parallel performance of DSMC simulations on unstructured grids, thus providing a feasible approach for efficient numerical simulation of high-Mach-number rarefied gas flows.
-
表 1 各模块可并行部分比例
Table 1. Proportion of parallelizable parts for each module
模块
名称比例
$ Q $/%核心可并行部分 子模块名称 子模块功能 初始化 92.6 Volume 网格面积计算 Subcell 子网格编号计算 Reassign 取样赋值初始化 粒子运动 23.6 Move 粒子坐标计算 粒子排序 94.1 Index 粒子排序 粒子碰撞 10.1 Colnsel 碰撞对数目计算 取样 100 Sample 网格信息取样 表 2 算法测试环境
Table 2. Algorithm testing environment
类别 名称 配置 硬件环境 CPU Intel(R) Core(TM)
i7-14700KFGPU NVIDIA GeForce
RTX 4070 SUPER内存 48 G DDR5 软件环境 操作系统 Ubuntu 22.04 编译器 NVFORTRAN 25.3 -
[1] RADTKE J, KEBSCHULL C, STOLL E. Interactions of the space debris environment with mega constellations: using the example of the OneWeb constellation[J]. Acta Astronautica, 2017, 131: 55-68. doi: 10.1016/j.actaastro.2016.11.021 [2] BIRD G A. Molecular gas dynamics and the direct simulation of gas flows[M]. Oxford: Oxford University Press, 1994. [3] 许啸, 王学德, 谭俊杰. 近空间高超声速流场通量分裂型DSMC-IP方法[J]. 航空动力学报, 2015, 30(2): 315-323. XU Xiao, WANG Xuede, TAN Junjie. Flux splitting DSMC-IP method in near space hypersonic flow field[J]. Journal of Aerospace Power, 2015, 30(2): 315-323. (in ChineseXU Xiao, WANG Xuede, TAN Junjie. Flux splitting DSMC-IP method in near space hypersonic flow field[J]. Journal of Aerospace Power, 2015, 30(2): 315-323. (in Chinese) [4] 王保国, 黄伟光, 钱耕, 等. 再入飞行中DSMC与Navier-Stokes两种模型的计算与分析[J]. 航空动力学报, 2011, 26(5): 961-976. WANG Baoguo, HUANG Weiguang, QIAN Geng, et al. Computation and analysis of DSMC and Navier-Stokes model in reentry flight[J]. Journal of Aerospace Power, 2011, 26(5): 961-976. (in ChineseWANG Baoguo, HUANG Weiguang, QIAN Geng, et al. Computation and analysis of DSMC and Navier-Stokes model in reentry flight[J]. Journal of Aerospace Power, 2011, 26(5): 961-976. (in Chinese) [5] 朱荣丽, 曹义华, 李栋, 等. 高超声速飞行器复杂流场过渡区DSMC数值模拟的一种新方案[J]. 宇航学报, 2006, 27(2): 167-171. ZHU Rongli, CAO Yihua, LI Dong, et al. A new version of hypersonic vehicle numerical simulation in direct simulation of Monte Carlo method of three dimensional flow in transitional regime[J]. Journal of Astronautics, 2006, 27(2): 167-171. (in ChineseZHU Rongli, CAO Yihua, LI Dong, et al. A new version of hypersonic vehicle numerical simulation in direct simulation of Monte Carlo method of three dimensional flow in transitional regime[J]. Journal of Astronautics, 2006, 27(2): 167-171. (in Chinese) [6] 李中华, 李志辉, 李海燕, 等. 过渡流区N-S/DSMC耦合计算研究[J]. 空气动力学学报, 2013, 31(3): 282-287. LI Zhonghua, LI Zhihui, LI Haiyan, et al. Study on N-S/DSMC coupling calculation in transition flow region[J]. Acta Aerodynamica Sinica, 2013, 31(3): 282-287. (in ChineseLI Zhonghua, LI Zhihui, LI Haiyan, et al. Study on N-S/DSMC coupling calculation in transition flow region[J]. Acta Aerodynamica Sinica, 2013, 31(3): 282-287. (in Chinese) [7] 王保国, 李学东, 刘淑艳. 高温高速稀薄流的DSMC算法与流场传热分析[J]. 航空动力学报, 2010, 25(6): 1203-1220. WANG Baoguo, LI Xuedong, LIU Shuyan. DSMC algorithm and heat transfer analysis of high temperature and high velocity rarefied gas flow[J]. Journal of Aerospace Power, 2010, 25(6): 1203-1220. (in ChineseWANG Baoguo, LI Xuedong, LIU Shuyan. DSMC algorithm and heat transfer analysis of high temperature and high velocity rarefied gas flow[J]. Journal of Aerospace Power, 2010, 25(6): 1203-1220. (in Chinese) [8] 梁杰, 阎超, 李志辉, 等. 稀薄过渡流区横向喷流干扰效应数值模拟研究[J]. 空气动力学学报, 2013, 31(1): 27-33. LIANG Jie, YAN Chao, LI Zhihui, et al. Numerical investigation of lateral jet interaction effects in rarefied transition flow regime[J]. Acta Aerodynamica Sinica, 2013, 31(1): 27-33. (in ChineseLIANG Jie, YAN Chao, LI Zhihui, et al. Numerical investigation of lateral jet interaction effects in rarefied transition flow regime[J]. Acta Aerodynamica Sinica, 2013, 31(1): 27-33. (in Chinese) [9] 王保国, 李耀华, 钱耕. 四种飞行器绕流的三维DSMC计算与传热分析[J]. 航空动力学报, 2011, 26(1): 1-20. WANG Baoguo, LI Yaohua, QIAN Geng. Three-dimensional DSMC calculation and heat transfer analysis of four capsules for hypersonic rarefied conditions[J]. Journal of Aerospace Power, 2011, 26(1): 1-20. (in ChineseWANG Baoguo, LI Yaohua, QIAN Geng. Three-dimensional DSMC calculation and heat transfer analysis of four capsules for hypersonic rarefied conditions[J]. Journal of Aerospace Power, 2011, 26(1): 1-20. (in Chinese) [10] 蔡国飙, 刘世俭, 王慧玉, 等. 真空羽流场的DSMC并行数值模拟[J]. 航空动力学报, 1999, 14(2): 113-118. CAI Guobiao, LIU Shijian, WANG Huiyu, et al. DSMC parallel numerical simulation of vacuum plume flow field[J]. Journal of Aerospace Power, 1999, 14(2): 113-118. (in ChineseCAI Guobiao, LIU Shijian, WANG Huiyu, et al. DSMC parallel numerical simulation of vacuum plume flow field[J]. Journal of Aerospace Power, 1999, 14(2): 113-118. (in Chinese) [11] 杜松蔚, 王学德. 一种基于双重空间网格结合的高效DSMC实现方法[J]. 航空动力学报, 2022, 37(9): 1846-1854. DU Songwei, WANG Xuede. High-efficiency DSMC implement method based on combination of dual spatial grids[J]. Journal of Aerospace Power, 2022, 37(9): 1846-1854. (in ChineseDU Songwei, WANG Xuede. High-efficiency DSMC implement method based on combination of dual spatial grids[J]. Journal of Aerospace Power, 2022, 37(9): 1846-1854. (in Chinese) [12] 张哲铭. 非结构网格DSMC的大规模并行计算研究[D]. 成都: 电子科技大学, 2022. ZHANG Zheming. Research on large-scale parallel computing of unstructured grid DSMC applications[D]. Chengdu: University of Electronic Science and Technology of China, 2022. (in ChineseZHANG Zheming. Research on large-scale parallel computing of unstructured grid DSMC applications[D]. Chengdu: University of Electronic Science and Technology of China, 2022. (in Chinese) [13] LUSK E, HUSS S, SAPHIR B, et al. MPI: a message-passing interface standard[J]. International Journal of Supercomputer Applications, 2009, 8(3/4): 623. [14] 王学德. 高超声速稀薄气流非结构网格DSMC及并行算法研究[D]. 南京: 南京航空航天大学, 2006. WANG Xuede. DSMC method on unstructured grids for hypersonic rarefied gas flow and its parallelization[D]. Nanjing: Nanjing University of Aeronautics and Astronautics, 2006. (in ChineseWANG Xuede. DSMC method on unstructured grids for hypersonic rarefied gas flow and its parallelization[D]. Nanjing: Nanjing University of Aeronautics and Astronautics, 2006. (in Chinese) [15] PassMark Software. Top CPU performance to date [EB/OL]. (2025-06-28)[2025-07-06]. https://www.cpubenchmark.net/year-on-year.html. [16] NVIDIA. NVIDIA CUDA toolkit documentation [EB/OL]. (2025-06-09) [2025-07-06]. https://developer.nvidia.com/cuda-faq. [17] CHEN Bojun. High performance parallel computing of DSMC method based on CUDA[EB/OL]. (2009-12-21) [2025-07-06]. https://wenku.baidu.com/view/f8f381ccbb4cf7ec4afed0cd.html?_wkts_=1763396744283 [18] 严立, 戴欣怡, 陈佳洛, 等. 基于计算统一设备架物Fortran的直接模拟蒙特卡洛方法并行优化[J]. 上海交通大学学报, 2013, 47(8): 1198-1204. YAN Li, DAI Xinyi, CHEN Jialuo, et al. Parallel optimization of direct simulation Monte Carlo method using compute unified device architecture fortran[J]. Journal of Shanghai Jiao Tong University, 2013, 47(8): 1198-1204. (in ChineseYAN Li, DAI Xinyi, CHEN Jialuo, et al. Parallel optimization of direct simulation Monte Carlo method using compute unified device architecture fortran[J]. Journal of Shanghai Jiao Tong University, 2013, 47(8): 1198-1204. (in Chinese) [19] 邱天林. 直角坐标网格下DSMC方法的GPU并行研究[D]. 南京: 南京航空航天大学, 2017. QIU Tianlin. Parallel research on GPU-based DSMC method using Cartesian grid[D]. Nanjing: Nanjing University of Aeronautics and Astronautics, 2017. (in ChineseQIU Tianlin. Parallel research on GPU-based DSMC method using Cartesian grid[D]. Nanjing: Nanjing University of Aeronautics and Astronautics, 2017. (in Chinese) [20] 陈飞同, 王学德. 三维复杂界面非结构网格N-S/DSMC耦合方法[J]. 航空动力学报, 2025, 40(9): 20240329. CHEN Feitong, WANG Xuede. N-S/DSMC coupling method using three-dimensional unstructured mesh for complex interfaces[J]. Journal of Aerospace Power, 2025, 40(9): 20240329. (in ChineseCHEN Feitong, WANG Xuede. N-S/DSMC coupling method using three-dimensional unstructured mesh for complex interfaces[J]. Journal of Aerospace Power, 2025, 40(9): 20240329. (in Chinese) [21] AMDAHL G M. Validity of the single processor approach to achieving large scale computing capabilities[C]//Proceedings of the Spring Joint Computer Conference. New York: Association for Computing Machinery, 1967: 483-485. [22] 王学德, 伍贻兆, 夏健, 等. 三维非结构网格DSMC并行算法及应用研究[J]. 宇航学报, 2007, 28(6): 1500-1505. WANG Xuede, WU Yizhao, XIA Jian, et al. A parallel algorithm of 3D unstructured DSMC method and its application[J]. Journal of Astronautics, 2007, 28(6): 1500-1505. (in ChineseWANG Xuede, WU Yizhao, XIA Jian, et al. A parallel algorithm of 3D unstructured DSMC method and its application[J]. Journal of Astronautics, 2007, 28(6): 1500-1505. (in Chinese) [23] 白洪涛. 基于GPU的高性能并行算法研究[D]. 长春: 吉林大学, 2010. BAI Hongtao. Research on high performance parallel algorithms based on GPU[D]. Changchun: Jilin University, 2010. (in ChineseBAI Hongtao. Research on high performance parallel algorithms based on GPU[D]. Changchun: Jilin University, 2010. (in Chinese) [24] 傅萌. 基于CPU+GPU并行计算加速的电力系统状态估计技术[D]. 南京: 东南大学, 2023. FU Meng. Power system state estimation technology based on CPU+GPU parallel computing acceleration[D]. Nanjing: Southeast University, 2023. (in ChineseFU Meng. Power system state estimation technology based on CPU+GPU parallel computing acceleration[D]. Nanjing: Southeast University, 2023. (in Chinese) [25] 鞠鹏飞, 宁方飞. GPU平台上的叶轮机械CFD加速计算[J]. 航空动力学报, 2014, 29(5): 1154-1162. JU Pengfei, NING Fangfei. Accelerated CFD computing of turbomachinery on GPU platform[J]. Journal of Aerospace Power, 2014, 29(5): 1154-1162. (in ChineseJU Pengfei, NING Fangfei. Accelerated CFD computing of turbomachinery on GPU platform[J]. Journal of Aerospace Power, 2014, 29(5): 1154-1162. (in Chinese) -

下载: