TensorFlow GPU环境执行fft2d运算触发InternalError报错
问题现象
- 相同逻辑的TensorFlow代码在CPU运行环境下执行无报错,最终输出张量形状为
TensorShape([2, 10, 20]) - Colab平台切换至GPU运行环境后,执行
tf.signal.fft2d二维快速傅里叶变换时抛出InternalError异常,计算中断
CPU环境可正常运行的复现代码
import tensorflow as tf sample_fft_input = tf.random.uniform((2, 10, 20)) sfi = tf.cast(sample_fft_input , tf.complex64) sfi = tf.math.real(tf.signal.fft2d(sfi)) print(sfi.shape) # 正常输出: TensorShape([2, 10, 20])
GPU环境报错信息
--------------------------------------------------------------------------- InternalError Traceback (most recent call last) <ipython-input-41-094e0f7d5037> in <module>() 3 sample_fft_input = tf.random.uniform((2, 10, 20)) 4 sfi = tf.cast(sample_fft_input, tf.complex64) ----> 5 sfi = tf.math.real(tf.signal.fft2d(sfi)) 6 print(sfi.shape) 1 frames /usr/local/lib/python3.7/dist-packages/tensorflow/python/framework/ops.py in raise_from_not_ok_status(e, name) 7162 def raise_from_not_ok_status(e, name): 7163 e.message += (" name: " + name if name is not None else "") -> 7164 raise core._status_to_exception(e) from None # pylint: disable=protected-access 7165 InternalError: fft failed : type=1 in.shape=[2,10,20] [Op:FFT2D]
故障原因与解决方法
这个报错是Colab预装的低版本TensorFlow调用GPU端cuFFT计算库时的适配bug导致的,CPU端FFT计算走的是FFTW库,不存在该兼容问题。可选修复方案如下:
- 强制指定FFT计算在CPU上运行,绕开GPU端的bug,代码改动最小:
import tensorflow as tf sample_fft_input = tf.random.uniform((2, 10, 20)) sfi = tf.cast(sample_fft_input , tf.complex64) with tf.device('/CPU:0'): sfi = tf.math.real(tf.signal.fft2d(sfi)) print(sfi.shape)
- 升级Colab内的TensorFlow到最新稳定版,修复旧版本的cuFFT调用缺陷。先在代码块执行
!pip install --upgrade tensorflow,之后重启Colab运行时再跑原有代码即可。 - 必须用GPU跑FFT的话,可以给输入张量最后两个维度补零到cuFFT适配性更好的长度(比如把维度10补到16、维度20补到32),FFT计算完成后再裁剪回原尺寸,也能避开该报错。
内容的提问来源于stack exchange,提问作者Libnist
相关产品推荐
相关产品推荐

