PyTorch GPU线性代数运算性能远逊于CPU的问题求助
问题:PyTorch GPU版线性代数运算速度远慢于CPU的优化建议
我正在对Numpy与PyTorch(CPU+GPU)进行基准测试,无法理解为何GPU版本的线性代数运算速度慢这么多。为避免CPU与GPU间的数据传输开销,计时仅针对线性代数运算部分,恳请各位提供优化建议。
测试代码
import torch import numpy as np import time import os os.environ["KMP_DUPLICATE_LIB_OK"]="TRUE" print(f"Is CUDA supported by this system? {torch.cuda.is_available()}") print(f"CUDA version: {torch.version.cuda}") # Storing ID of current CUDA device cuda_id = torch.cuda.current_device() print(f"ID of current CUDA device: {torch.cuda.current_device()}") print(f"Name of current CUDA device: {torch.cuda.get_device_name(cuda_id)}") device = torch.device("cuda" if torch.cuda.is_available() else "cpu") size=10000 real=1 A=np.random.rand(size,size) b=np.random.rand(size,1) start_time = time.time() for t in range(real): x_np=np.linalg.solve(A,b) print("NUMPY CPU--- %s seconds ---" % (time.time() - start_time)) A = torch.from_numpy(A) b = torch.from_numpy(b) start_time = time.time() for t in range(real): x_torch = torch.linalg.solve(A, b) print("PYTORCH CPU--- %s seconds ---" % (time.time() - start_time)) A = A.to(device) b = b.to(device) start_time = time.time() for t in range(real): x_torch = torch.linalg.solve(A, b) print("PYTORCH GPU--- %s seconds ---" % (time.time() - start_time))
测试结果
本系统是否支持CUDA? True CUDA版本: 12.1 当前CUDA设备ID: 0 当前CUDA设备名称: NVIDIA GeForce RTX 3060 NUMPY CPU--- 1.6754064559936523 秒 --- PYTORCH CPU--- 1.3463587760925293 秒 --- PYTORCH GPU--- 3.8940138816833496 秒 ---
原因分析与优化建议
1. GPU异步操作导致计时偏差
PyTorch的CUDA操作默认异步执行,time.time()仅记录任务提交时间,而非GPU实际完成运算的时间。必须通过torch.cuda.synchronize()强制同步GPU,确保时间统计准确。
2. GPU内核启动开销占比过高
单次运算时,GPU的内核加载、初始化等固定开销占比极大,掩盖了并行计算的优势。通过预热运算和增加重复次数,可以分摊这部分开销,体现GPU的真实性能。
3. 未启用硬件加速优化
默认配置下,PyTorch可能未充分利用CuDNN等硬件加速库的最优算法。开启torch.backends.cudnn.benchmark = True可让系统自动选择适配当前硬件和输入规模的最优计算路径。
4. 矩阵类型与规模适配问题
GPU对正定矩阵等特定类型的线性代数运算有专门优化,可尝试生成正定矩阵测试;同时,更大的矩阵规模或多次重复运算更能发挥GPU的并行优势。
优化后的测试代码
import torch import numpy as np import time import os os.environ["KMP_DUPLICATE_LIB_OK"]="TRUE" print(f"本系统是否支持CUDA? {torch.cuda.is_available()}") print(f"CUDA版本: {torch.version.cuda}") cuda_id = torch.cuda.current_device() print(f"当前CUDA设备ID: {torch.cuda.current_device()}") print(f"当前CUDA设备名称: {torch.cuda.get_device_name(cuda_id)}") device = torch.device("cuda" if torch.cuda.is_available() else "cpu") # 开启CuDNN基准测试,自动选择最优算法 torch.backends.cudnn.benchmark = True size=10000 real=10 # 增加重复次数,分摊启动开销 # 生成正定矩阵,适配GPU优化 A = np.random.rand(size,size) A = A @ A.T + np.eye(size) * 1e-3 # 保证矩阵正定 b=np.random.rand(size,1) # Numpy CPU测试 start_time = time.time() for t in range(real): x_np=np.linalg.solve(A,b) print(f"NUMPY CPU--- {time.time() - start_time:.4f} 秒 ---") # PyTorch CPU测试 A_torch = torch.from_numpy(A) b_torch = torch.from_numpy(b) start_time = time.time() for t in range(real): x_torch_cpu = torch.linalg.solve(A_torch, b_torch) print(f"PYTORCH CPU--- {time.time() - start_time:.4f} 秒 ---") # PyTorch GPU测试:预热+同步 A_gpu = A_torch.to(device) b_gpu = b_torch.to(device) # 预热运算,加载GPU内核 for _ in range(2): torch.linalg.solve(A_gpu, b_gpu) torch.cuda.synchronize() # 等待预热完成 # 正式计时,确保GPU同步 start_time = time.time() torch.cuda.synchronize() for t in range(real): x_torch_gpu = torch.linalg.solve(A_gpu, b_gpu) torch.cuda.synchronize() # 等待所有GPU运算完成 print(f"PYTORCH GPU--- {time.time() - start_time:.4f} 秒 ---")
内容的提问来源于stack exchange,提问作者Marios Karaoulis
相关产品推荐
相关产品推荐

