GPU-Accelerated Implementation of Two-Layer Boussinesq Model for Fluids Physics
A team of researchers from Dalian University of Technology has developed a new implementation of a high-accuracy two-layer Boussinesq model for fluids physics, utilizing graphics processing units (GPUs) to achieve substantial computational speedups. The model, funded by the National Natural Science Foundation of China, has been extensively optimized to align with the new parallel computing paradigm, resulting in improved performance.
Key Takeaways:
- The researchers presented a GPU-accelerated implementation of a high-accuracy two-layer Boussinesq model using Compute Unified Device Architecture (CUDA).
- The solution procedure of the original model has been extensively optimized to align with the new parallel computing paradigm, resulting in substantial computational speedups.
- The optimizations include the restructuring of the solution sequence for the two-layer vertical velocity, reorganization and solving in parallel blocks, and refinement of the tolerance calculation module through atomic operations.
- The cyclic reduction algorithm has been extended using ghost grids and the concatenation matrix method, with three optimization strategies designed for different variables.
- A series of three-dimensional simulations were conducted to validate the accuracy of the code, the precision of the simulations, and computational efficiency, including deep-water wave group propagation and nearshore wave evolution.
- The comparison results show that the model can accurately capture the nonlinearity and dispersion of waves from deep to shallow water.
- For single-GPU computation, the model achieved 4.11x and 2.5x speedup compared to serial and parallel central processing unit (CPU) code in larger computing grids (2.6x10^5).
- The research has been peer-reviewed and published in the journal Physics of Fluids.
Statistics:
- 4.11x speedup achieved for single-GPU computation compared to serial CPU code.
- 2.5x speedup achieved for single-GPU computation compared to parallel CPU code.
- Computational efficiency improved by reconstructing the solution sequence for the two-layer vertical velocity and refining the tolerance calculation module.
- The cyclic reduction algorithm was extended using ghost grids and the concatenation matrix method, resulting in additional optimization strategies.
Sources:
- A Two-layer Boussinesq-type Wave Model Accelerated By Graphics Processing Unit. Physics of Fluids, 2025;37(7).
- Dalian University of Technology, State Key Laboratory of Coastal and Offshore Engineering, Dalian 116024, People's Republic of China.
- Kezhao Fang, Dalian University of Technology.
- Guanglin Chen, Dalian University of Technology.
- Jiawen Sun, Dalian University of Technology.
- Ping Wang, Dalian University of Technology.
- Zhongbo Liu, Dalian University of Technology.
- Hong Jin, Dalian University of Technology.
- National Natural Science Foundation of China (NSFC).
- Journal of Physics Research, August 5, 2025; p 194.