You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyTorch GPU线性代数运算性能远逊于CPU的问题求助

问题:PyTorch GPU版线性代数运算速度远慢于CPU的优化建议

我正在对Numpy与PyTorch(CPU+GPU)进行基准测试,无法理解为何GPU版本的线性代数运算速度慢这么多。为避免CPU与GPU间的数据传输开销,计时仅针对线性代数运算部分,恳请各位提供优化建议。

测试代码

import torch
import numpy as np
import time
import os

os.environ["KMP_DUPLICATE_LIB_OK"]="TRUE"

print(f"Is CUDA supported by this system?  {torch.cuda.is_available()}")
print(f"CUDA version: {torch.version.cuda}")

# Storing ID of current CUDA device
cuda_id = torch.cuda.current_device()
print(f"ID of current CUDA device: {torch.cuda.current_device()}")
   
print(f"Name of current CUDA device: {torch.cuda.get_device_name(cuda_id)}")
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")


size=10000
real=1
A=np.random.rand(size,size)
b=np.random.rand(size,1)
start_time = time.time()
for t in range(real):
    x_np=np.linalg.solve(A,b)
print("NUMPY CPU--- %s seconds ---" % (time.time() - start_time))

A = torch.from_numpy(A)
b = torch.from_numpy(b)
start_time = time.time()
for t in range(real):
    x_torch = torch.linalg.solve(A, b)
print("PYTORCH CPU--- %s seconds ---" % (time.time() - start_time))

A = A.to(device)
b = b.to(device)
start_time = time.time()
for t in range(real):
    x_torch = torch.linalg.solve(A, b)
print("PYTORCH GPU--- %s seconds ---" % (time.time() - start_time))

测试结果

本系统是否支持CUDA?  True
CUDA版本: 12.1
当前CUDA设备ID: 0
当前CUDA设备名称: NVIDIA GeForce RTX 3060
NUMPY CPU--- 1.6754064559936523 秒 ---
PYTORCH CPU--- 1.3463587760925293 秒 ---
PYTORCH GPU--- 3.8940138816833496 秒 ---

原因分析与优化建议

1. GPU异步操作导致计时偏差

PyTorch的CUDA操作默认异步执行,time.time()仅记录任务提交时间,而非GPU实际完成运算的时间。必须通过torch.cuda.synchronize()强制同步GPU,确保时间统计准确。

2. GPU内核启动开销占比过高

单次运算时,GPU的内核加载、初始化等固定开销占比极大,掩盖了并行计算的优势。通过预热运算和增加重复次数,可以分摊这部分开销,体现GPU的真实性能。

3. 未启用硬件加速优化

默认配置下,PyTorch可能未充分利用CuDNN等硬件加速库的最优算法。开启torch.backends.cudnn.benchmark = True可让系统自动选择适配当前硬件和输入规模的最优计算路径。

4. 矩阵类型与规模适配问题

GPU对正定矩阵等特定类型的线性代数运算有专门优化,可尝试生成正定矩阵测试;同时,更大的矩阵规模或多次重复运算更能发挥GPU的并行优势。

优化后的测试代码

import torch
import numpy as np
import time
import os

os.environ["KMP_DUPLICATE_LIB_OK"]="TRUE"

print(f"本系统是否支持CUDA?  {torch.cuda.is_available()}")
print(f"CUDA版本: {torch.version.cuda}")

cuda_id = torch.cuda.current_device()
print(f"当前CUDA设备ID: {torch.cuda.current_device()}")
print(f"当前CUDA设备名称: {torch.cuda.get_device_name(cuda_id)}")
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")

# 开启CuDNN基准测试,自动选择最优算法
torch.backends.cudnn.benchmark = True

size=10000
real=10  # 增加重复次数,分摊启动开销
# 生成正定矩阵,适配GPU优化
A = np.random.rand(size,size)
A = A @ A.T + np.eye(size) * 1e-3  # 保证矩阵正定
b=np.random.rand(size,1)

# Numpy CPU测试
start_time = time.time()
for t in range(real):
    x_np=np.linalg.solve(A,b)
print(f"NUMPY CPU--- {time.time() - start_time:.4f} 秒 ---")

# PyTorch CPU测试
A_torch = torch.from_numpy(A)
b_torch = torch.from_numpy(b)
start_time = time.time()
for t in range(real):
    x_torch_cpu = torch.linalg.solve(A_torch, b_torch)
print(f"PYTORCH CPU--- {time.time() - start_time:.4f} 秒 ---")

# PyTorch GPU测试:预热+同步
A_gpu = A_torch.to(device)
b_gpu = b_torch.to(device)

# 预热运算,加载GPU内核
for _ in range(2):
    torch.linalg.solve(A_gpu, b_gpu)
torch.cuda.synchronize()  # 等待预热完成

# 正式计时,确保GPU同步
start_time = time.time()
torch.cuda.synchronize()
for t in range(real):
    x_torch_gpu = torch.linalg.solve(A_gpu, b_gpu)
torch.cuda.synchronize()  # 等待所有GPU运算完成
print(f"PYTORCH GPU--- {time.time() - start_time:.4f} 秒 ---")

内容的提问来源于stack exchange,提问作者Marios Karaoulis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 13:45:28