You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

重复单精度复数矩阵向量乘法:性能与精度优化问询

复数矩阵向量批量乘法的性能与精度优化问题

我将一个耗时函数简化为一系列重复的单精度复数矩阵向量乘法,矩阵固定但待处理向量数量庞大,测试程序如下:

module maths

contains
subroutine lots_of_MVM(Y,P,t1,t2,startRow)
    implicit none
    ! args
    complex, intent(in), contiguous    :: Y(:,:),P(:,:)
    complex, intent(inout), contiguous :: t1(:,:),t2(:,:)
    integer, intent(in)                :: startRow
    
    ! locals
    integer :: ii,jj,zz,nrhs,n,pCol,tRow,yCol
    ! indexing
    nrhs = size(P,2)/2
    n    = size(Y,1)
    
    ! Do lots of maths
    !$OMP PARALLEL PRIVATE(jj,pCol,tRow,yCol,zz)
    !$OMP DO
    do jj=1,nrhs
        pCol = jj*2-1
        tRow = startRow
        do yCol=1,size(Y,2)
            ! This is faster than doing sum(P(:,pCol)*Y(:,yCol))
            do zz=1,n
                t1(tRow,jj) = t1(tRow,jj) + P(zz,pCol  )*Y(zz,yCol)
                t2(tRow,jj) = t2(tRow,jj) + P(zz,pCol+1)*Y(zz,yCol)
            end do
            tRow = tRow + 1
        end do
    end do
    !$OMP END DO
    !$OMP END PARALLEL
    
end subroutine

end module
    
program test
    use maths
    use omp_lib
    implicit none
    ! variables
    complex, allocatable,dimension(:,:) :: Y,P,t1,t2
    integer :: n,nrhs,nY,yStart,yStop,mult
    double precision startTime
    
    ! setup (change mult to make problem larger)
    ! real problem size mult = 1000 to 2000
    mult = 300
    n = 10*mult
    nY = 30*mult
    nrhs = 20*mult
    yStart = 5
    yStop  = yStart + nrhs - 1
    
    ! allocate
    allocate(Y(n,nY),P(n,nrhs*2))
    allocate(t1(nrhs,nrhs),t2(nrhs,nrhs))
    
    ! make some data
    call random_number(Y%re)
    call random_number(Y%im)
    call random_number(P%re) 
    call random_number(P%im)
    t1 = 0
    t2 = 0
    
    ! do maths
    startTime = omp_get_wtime()
    call lots_of_MVM(Y(:,yStart:yStop),P,t1,t2,1)
    write(*,*) omp_get_wtime()-startTime
end program

我已尝试多种性能优化方案,当前实现为最优版本:

  • 64字节边界数据对齐:无明显效果,且不希望增加内存开销
  • MKL cgemv_batch_strided:性能不如当前OMP嵌套循环实现
  • 交换循环顺序:需后续执行转置操作,未采用

目前的核心需求是在不显著降低性能的前提下提升计算精度,但已尝试的精度优化方案均导致性能大幅下降。

限制条件

  • 仅可使用2-4个OMP核心
  • 不可大幅增加内存消耗

其他环境与背景信息

  • 运行环境:RHEL 8系统,使用Intel Fortran编译器,编译选项为-O3 -xHost
  • 测试规模:可将mult设为1000-2000以模拟真实问题量级
  • 硬件平台:当前运行于双路Intel系统,未来可能迁移至AMD平台
  • 多实例部署需求:当前双路28核768GB内存系统可同时运行24个实例(每个实例占用2核)
  • 代码定位:此模块为大型代码中最耗时部分,其余模块多为单线程,难以并行化
  • GPU版本情况:已实现CUDA batched cgemv GPU版本,单实例性能比CPU版本快6倍,但GPU单实例会饱和带宽/算力,CPU平台的总吞吐量更高

内容的提问来源于stack exchange,提问作者js1

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 23:27:12