Fortran派生类型数组求和赋值性能偏低是否正常?如何优化?
Fortran派生类型与原生real数组的性能对比及优化建议
测试场景
对比Fortran原生real(8)类型二维数组,与仅包含单个real(8)二维数组的派生类型,执行c = a + b形式的求和赋值操作的性能差异。
代码实现
派生类型定义模块
module type_mod use iso_fortran_env type :: class_t real(8), dimension(:,:), allocatable :: a contains procedure :: assign_type generic, public :: assignment(=) => assign_type procedure :: sum_type generic :: operator(+) => sum_type final :: destroy end type class_t contains subroutine assign_type(lhs, rhs) class(class_t), intent(inout) :: lhs type(class_t), intent(in) :: rhs lhs % a = rhs % a end subroutine assign_type subroutine destroy(this) type(class_t), intent(inout) :: this if (allocated(this % a)) deallocate(this % a) end subroutine destroy function sum_type (lhs, rhs) result(res) class(class_t), intent(in) :: lhs type(class_t), intent(in) :: rhs type(class_t) :: res res % a = lhs % a + rhs % a end function sum_type end module type_mod
性能对比子例程模块
module subroutine_mod use type_mod, only: class_t contains subroutine sum_real(a, b, c) real(8), dimension(:,:), intent(inout) :: a, b, c c = a + b end subroutine sum_real subroutine sum_type(a, b, c) type(class_t), intent(inout) :: a, b, c c = a + b end subroutine sum_type end module subroutine_mod
测试主程序
测试使用尺寸为(10000,10000)的数组,重复执行求和赋值操作100次:
program test use subroutine_mod integer :: i integer :: N = 100 ! Number of times to repeat the assign integer :: M = 10000 ! Size of the arrays real(8) :: tf, ts real(8), dimension(:,:), allocatable :: a, b, c type(class_t) :: a2, b2, c2 allocate(a2%a(M,M), b2%a(M,M), c2%a(M,M)) a2%a = 1.0d0 b2%a = 2.0d0 c2%a = 3.0d0 allocate(a(M,M), b(M,M), c(M,M)) a = 1.0d0 b = 2.0d0 c = 3.0d0 ! Benchmark timing with call cpu_time(ts) do i = 1, N call sum_type(a2, b2, c2) end do call cpu_time(tf) write(*,*) "Type : ", tf-ts call cpu_time(ts) do i = 1, N call sum_real(a, b, c) end do call cpu_time(tf) write(*,*) "Real : ", tf-ts end program test
测试结果
使用-O2优化编译后,两种编译器下均出现明显性能差异:
gfortran 9.4.0
- 派生类型操作:33秒
- 原生real数组操作:13秒
ifort 2021.6.0
- 派生类型操作:30秒
- 原生real数组操作:3秒
核心问题
这种性能差异是否属于正常现象?如果是,有哪些性能优化建议?
背景说明
该仅包含单个数组的派生类型用于代码重构,需要与原有代码保持相似接口。
性能差异是否正常?
这种性能差异是正常现象,主要原因有两点:
- 临时对象的额外开销:派生类型的
operator(+)函数会创建临时class_t对象,c = a + b的执行流程是:先调用sum_type生成包含求和结果的临时派生对象,再调用assign_type将临时对象的数组赋值给c%a,最后销毁临时对象。这一过程比原生数组直接的c = a + b多了临时对象的创建、赋值、销毁步骤,带来额外的内存操作和函数调用开销。 - 编译器优化能力差异:编译器对原生数组的优化非常成熟,可直接生成向量指令、循环展开等高效代码;而派生类型的封装层会阻碍编译器的优化穿透,尤其是旧版本编译器对派生类型过程的内联优化支持不足,导致底层数组操作无法被充分优化。
性能优化建议
- 强制过程内联:在派生类型的过程定义中添加
inline属性,让编译器将小过程直接嵌入调用点,消除函数调用开销并暴露底层数组操作以进行优化。例如:type :: class_t real(8), dimension(:,:), allocatable :: a contains procedure, inline :: assign_type generic, public :: assignment(=) => assign_type procedure, inline :: sum_type generic :: operator(+) => sum_type final :: destroy end type class_t - 避免不必要的临时对象:若业务场景允许,可重载
operator(+=)实现原地求和,或直接定义三参数求和子例程(如subroutine add(lhs, rhs, res)),跳过临时对象的创建过程,直接将结果写入目标数组。 - 添加
contiguous属性:在派生类型的数组定义中添加contiguous属性,明确数组内存连续性,帮助编译器优化内存访问模式:real(8), dimension(:,:), allocatable, contiguous :: a - 升级编译器版本:旧版本编译器对派生类型的优化支持有限,升级到gfortran 13+或Intel oneAPI Fortran Compiler等新版本,可获得更好的派生类型优化能力。
- 权衡final过程的必要性:若临时对象创建频繁,final过程的销毁操作会带来额外开销。如果代码可通过其他方式安全管理内存,可考虑移除final过程以提升性能。
内容的提问来源于stack exchange,提问作者Antônio Carneiro
相关产品推荐
相关产品推荐

