如何在nvfortran编译器中启用内存越界检查?
描述及示例代码
以下是两个简单的测试程序,分别在CPU和GPU代码中尝试访问越界内存。我将GPU示例单独列出,以便可以使用不同编译器测试CPU示例并观察其行为。
CPU示例
module sizes integer, save :: size1 integer, save :: size2 end module sizes module arrays real, allocatable, save :: testArray1(:, :) real, allocatable, save :: testArray2(:, :) end module arrays subroutine testMemoryAccess use sizes use arrays implicit none real :: value value = testArray1(size1+1, size2+1) print *, 'value', value end subroutine testMemoryAccess Program testMemoryAccessOutOfBounds use sizes use arrays implicit none ! set sizes for the example size1 = 5000 size2 = 2500 allocate (testArray1(size1, size2)) allocate (testArray2(size2, size1)) testArray1 = 1.d0 testArray2 = 2.d0 call testMemoryAccess end program testMemoryAccessOutOfBounds
GPU示例
module sizes integer, save :: size1 integer, save :: size2 end module sizes module sizesCuda integer, device, save :: size1 integer, device, save :: size2 end module sizesCuda module arrays real, allocatable, save :: testArray1(:, :) real, allocatable, save :: testArray2(:, :) end module arrays module arraysCuda real, allocatable, device, save :: testArray1(:, :) real, allocatable, device, save :: testArray2(:, :) end module arraysCuda module cudaKernels use cudafor use sizesCuda use arraysCuda contains attributes(global) Subroutine testMemoryAccessCuda implicit none integer :: element real :: value element = (blockIdx%x - 1)*blockDim%x + threadIdx%x if (element.eq.1) then value = testArray1(size1+1, size2+1) print *, 'value', value end if end Subroutine testMemoryAccessCuda end module cudaKernels Program testMemoryAccessOutOfBounds use cudafor use cudaKernels use sizes use sizesCuda, size1_d => size1, size2_d => size2 use arrays use arraysCuda, testArray1_d => testArray1, testArray2_d => testArray2 implicit none integer :: istat ! set sizes for the example size1 = 5000 size2 = 2500 size1_d = size1 size2_d = size2 allocate (testArray1_d(size1, size2)) allocate (testArray2_d(size2, size1)) testArray1_d = 1.d0 testArray2_d = 2.d0 call testMemoryAccessCuda<<<64, 64>>> istat = cudadevicesynchronize() end program testMemoryAccessOutOfBounds
使用nvfortran调试程序时,编译器未对越界访问给出任何警告。查看可用的越界访问检测标志,-C和-Mbounds选项似乎应该实现该功能,但实际未按预期工作。
使用ifort进行相同测试时,编译器会停止运行并打印出越界访问发生的具体行号。
如何通过nvfortran实现这一功能?我原本以为这是CUDA特有的问题,但在编写示例代码时发现,nvfortran在CPU代码中也存在同样的问题,因此这并非CUDA相关问题。
使用的编译器
nvfortran
nvfortran 23.5-0 64-bit target on x86-64 Linux -tp zen2 NVIDIA Compilers and Tools Copyright (c) 2022, NVIDIA CORPORATION & AFFILIATES. All rights reserved.
ifort
ifort (IFORT) 2021.10.0 20230609 Copyright (C) 1985-2023 Intel Corporation. All rights reserved.
操作步骤
nvfortran
我按以下方式编译示例:
nvfortran -C -traceback -Mlarge_arrays -Mdclchk -cuda -gpu=cc86 testOutOfBounds.f90
nvfortran -C -traceback -Mlarge_arrays -Mdclchk -cuda -gpu=cc86 testOutOfBoundsCuda.f90
运行CPU代码时,得到未初始化的数组值:
value 1.5242136E-27
运行GPU代码时,得到零值:
value 0.000000
ifort
我按以下方式编译CPU示例:
ifort -init=snan -C -fpe0 -g -traceback testOutOfBounds.f90
得到如下输出:
forrtl: severe (408): fort: (2): Subscript #2 of the array TESTARRAY1 has value 2501 which is greater than the upper bound of 2500 Image PC Routine Line Source a.out 00000000004043D4 testmemoryaccess_ 23 testOutOfBounds.f90 a.out 0000000000404FD6 MAIN__ 43 testOutOfBounds.f90 a.out 000000000040418D Unknown Unknown Unknown libc.so.6 00007F65A9229D90 Unknown Unknown Unknown libc.so.6 00007F65A9229E40 __libc_start_main Unknown Unknown a.out 00000000004040A5 Unknown Unknown Unknown
这正是我期望编译器输出的结果。
内容的提问来源于stack exchange,提问作者SadBoySquad
相关产品推荐
相关产品推荐

