You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Fortran派生类型数组求和赋值性能偏低是否正常?如何优化?

Fortran派生类型与原生real数组的性能对比及优化建议

测试场景

对比Fortran原生real(8)类型二维数组,与仅包含单个real(8)二维数组的派生类型,执行c = a + b形式的求和赋值操作的性能差异。

代码实现

派生类型定义模块

module type_mod

use iso_fortran_env

type :: class_t
  real(8), dimension(:,:), allocatable :: a
contains
  procedure :: assign_type
  generic, public :: assignment(=) => assign_type
  procedure :: sum_type
  generic :: operator(+) => sum_type
  final :: destroy
end type class_t

contains

  subroutine assign_type(lhs, rhs)
    class(class_t), intent(inout) :: lhs
    type(class_t), intent(in) :: rhs
    lhs % a = rhs % a
  end subroutine assign_type

  subroutine destroy(this)
    type(class_t), intent(inout) :: this
    if (allocated(this % a)) deallocate(this % a)
  end subroutine destroy

  function sum_type (lhs, rhs) result(res)
    class(class_t), intent(in) :: lhs
    type(class_t), intent(in) :: rhs
    type(class_t) :: res
    res % a = lhs % a + rhs % a
  end function sum_type

end module type_mod

性能对比子例程模块

module subroutine_mod

  use type_mod, only: class_t

  contains

  subroutine sum_real(a, b, c)
    real(8), dimension(:,:), intent(inout) :: a, b, c
    c = a + b
  end subroutine sum_real


  subroutine sum_type(a, b, c)
    type(class_t), intent(inout) :: a, b, c
    c = a + b
  end subroutine sum_type

end module subroutine_mod

测试主程序

测试使用尺寸为(10000,10000)的数组,重复执行求和赋值操作100次:

program test

  use subroutine_mod

  integer :: i
  integer :: N = 100 ! Number of times to repeat the assign
  integer :: M = 10000 ! Size of the arrays
  real(8) :: tf, ts
  real(8), dimension(:,:), allocatable :: a, b, c
  type(class_t) :: a2, b2, c2

  allocate(a2%a(M,M), b2%a(M,M), c2%a(M,M))
  a2%a = 1.0d0
  b2%a = 2.0d0
  c2%a = 3.0d0

  allocate(a(M,M), b(M,M), c(M,M))
  a = 1.0d0
  b = 2.0d0
  c = 3.0d0

  ! Benchmark timing with
  call cpu_time(ts)
  do i = 1, N
    call sum_type(a2, b2, c2)
  end do
  call cpu_time(tf)
  write(*,*) "Type : ", tf-ts

  call cpu_time(ts)
  do i = 1, N
    call sum_real(a, b, c)
  end do
  call cpu_time(tf)
  write(*,*) "Real : ", tf-ts

end program test

测试结果

使用-O2优化编译后,两种编译器下均出现明显性能差异:

gfortran 9.4.0

  • 派生类型操作:33秒
  • 原生real数组操作:13秒

ifort 2021.6.0

  • 派生类型操作:30秒
  • 原生real数组操作:3秒

核心问题

这种性能差异是否属于正常现象?如果是,有哪些性能优化建议?

背景说明

该仅包含单个数组的派生类型用于代码重构,需要与原有代码保持相似接口。


性能差异是否正常?

这种性能差异是正常现象,主要原因有两点:

  1. 临时对象的额外开销:派生类型的operator(+)函数会创建临时class_t对象,c = a + b的执行流程是:先调用sum_type生成包含求和结果的临时派生对象,再调用assign_type将临时对象的数组赋值给c%a,最后销毁临时对象。这一过程比原生数组直接的c = a + b多了临时对象的创建、赋值、销毁步骤,带来额外的内存操作和函数调用开销。
  2. 编译器优化能力差异:编译器对原生数组的优化非常成熟,可直接生成向量指令、循环展开等高效代码;而派生类型的封装层会阻碍编译器的优化穿透,尤其是旧版本编译器对派生类型过程的内联优化支持不足,导致底层数组操作无法被充分优化。

性能优化建议

  • 强制过程内联:在派生类型的过程定义中添加inline属性,让编译器将小过程直接嵌入调用点,消除函数调用开销并暴露底层数组操作以进行优化。例如:
    type :: class_t
      real(8), dimension(:,:), allocatable :: a
    contains
      procedure, inline :: assign_type
      generic, public :: assignment(=) => assign_type
      procedure, inline :: sum_type
      generic :: operator(+) => sum_type
      final :: destroy
    end type class_t
    
  • 避免不必要的临时对象:若业务场景允许,可重载operator(+=)实现原地求和,或直接定义三参数求和子例程(如subroutine add(lhs, rhs, res)),跳过临时对象的创建过程,直接将结果写入目标数组。
  • 添加contiguous属性:在派生类型的数组定义中添加contiguous属性,明确数组内存连续性,帮助编译器优化内存访问模式:
    real(8), dimension(:,:), allocatable, contiguous :: a
    
  • 升级编译器版本:旧版本编译器对派生类型的优化支持有限,升级到gfortran 13+或Intel oneAPI Fortran Compiler等新版本,可获得更好的派生类型优化能力。
  • 权衡final过程的必要性:若临时对象创建频繁,final过程的销毁操作会带来额外开销。如果代码可通过其他方式安全管理内存,可考虑移除final过程以提升性能。

内容的提问来源于stack exchange,提问作者Antônio Carneiro

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 15:50:39