You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyTorch Lightning启用GPU训练时CUBLAS_STATUS_INVALID_VALUE错误求助

问题描述

我正在学习PyTorch Lightning官方教程,为尝试GPU训练,将trainer的定义修改为:

trainer = pl.Trainer(limit_train_batches=100, max_epochs=1, gpus=1)

之后出现以下错误:

RuntimeError                              Traceback (most recent call last)
Cell In [3], line 4
      1 # train the model (hint: here are some helpful Trainer arguments for rapid idea iteration)
      2 # trainer = pl.Trainer(limit_train_batches=100, max_epochs=3)
      3 trainer = pl.Trainer(limit_train_batches=100, max_epochs=3, accelerator='gpu', devices=1)
----> 4 trainer.fit(model=autoencoder, train_dataloaders=train_loader)

File ~/miniconda3/envs/py38-cu116/lib/python3.8/site-packages/pytorch_lightning/trainer/trainer.py:696, in Trainer.fit(self, model, train_dataloaders, val_dataloaders, datamodule, ckpt_path)
    677 """
    678 Runs the full optimization routine.
    679 
   (...)
    693     datamodule: An instance of :class:`~pytorch_lightning.core.datamodule.LightningDataModule`.
    694 """
    695 self.strategy.model = model
---> 696 self._call_and_handle_interrupt(
    697     self._fit_impl, model, train_dataloaders, val_dataloaders, datamodule, ckpt_path
    698 )

File ~/miniconda3/envs/py38-cu116/lib/python3.8/site-packages/pytorch_lightning/trainer/trainer.py:650, in Trainer._call_and_handle_interrupt(self, trainer_fn, *args, **kwargs)
    648         return self.strategy.launcher.launch(trainer_fn, *args, trainer=self, **kwargs)
    649     else:
---> 650         return trainer_fn(*args, **kwargs)
    651 # TODO(awaelchli): Unify both exceptions below, where `KeyboardError` doesn't re-raise
    652 except KeyboardInterrupt as exception:

[...]

File ~/miniconda3/envs/py38-cu116/lib/python3.8/site-packages/pytorch_lightning/core/module.py:1450, in LightningModule.backward(self, loss, optimizer, optimizer_idx, *args, **kwargs)
   1433 def backward(
   1434     self, loss: Tensor, optimizer: Optional[Optimizer], optimizer_idx: Optional[int], *args, **kwargs
   1435 ) -> None:
   1436     """Called to perform backward on the loss returned in :meth:`training_step`. Override this hook with your
   1437     own implementation if you need to.
   1438 
   (...)
   1448             loss.backward()
   1449     """
-> 1450     loss.backward(*args, **kwargs)

File ~/miniconda3/envs/py38-cu116/lib/python3.8/site-packages/torch/_tensor.py:396, in Tensor.backward(self, gradient, retain_graph, create_graph, inputs)
    387 if has_torch_function_unary(self):
    388     return handle_torch_function(
    389         Tensor.backward,
    390         (self,),
   (...)
    394         create_graph=create_graph,
    395         inputs=inputs)
---> 396 torch.autograd.backward(self, gradient, retain_graph, create_graph, inputs=inputs)

File ~/miniconda3/envs/py38-cu116/lib/python3.8/site-packages/torch/autograd/__init__.py:173, in backward(tensors, grad_tensors, retain_graph, create_graph, grad_variables, inputs)
    168     retain_graph = create_graph
    170 # The reason we repeat same the comment below is that
    171 # some Python versions print out the first line of a multi-line function
    172 # calls in the traceback and some print out the last line
---> 173 Variable._execution_engine.run_backward(  # Calls into the C++ engine to run the backward pass
    174     tensors, grad_tensors_, retain_graph, create_graph, inputs,
    175     allow_unreachable=True, accumulate_grad=True)

RuntimeError: CUDA error: CUBLAS_STATUS_INVALID_VALUE when calling `cublasSgemm( handle, opa, opb, m, n, k, &alpha, a, lda, b, ldb, &beta, c, ldc)`

尝试使用devices=1, accelerator='ddp'替代时,出现错误:

ValueError: You selected an invalid accelerator name: `accelerator='ddp'`. Available names are: cpu, cuda, hpu, ipu, mps, tpu.

环境信息:

  • CUDA 11.6
  • Python 3.8.13
  • PyTorch 1.12.1
  • PyTorch Lightning 1.7.7
解决方法

1. 修正加速器参数错误

ddp是分布式训练策略,不属于加速器类型。在PyTorch Lightning 1.7.x版本中,正确的GPU训练配置有两种方式:

# 方式1:使用新版本推荐的参数
trainer = pl.Trainer(limit_train_batches=100, max_epochs=1, accelerator="cuda", devices=1)

# 方式2:兼容旧版本的gpus参数(1.7.x仍支持)
trainer = pl.Trainer(limit_train_batches=100, max_epochs=1, gpus=1)

2. 解决CUBLAS_STATUS_INVALID_VALUE错误

这个错误通常和张量维度不匹配、数据类型不一致或CUDA环境问题有关,按以下步骤排查:

  • 检查张量维度:确保模型前向传播、损失计算中的所有张量维度匹配,比如自动编码器的输入和输出维度是否一致,矩阵乘法的维度是否兼容。
  • 统一数据类型:确认模型参数和输入数据为同一类型(比如均为float32,避免混合float16和float32)。可以在数据加载时显式转换类型,或在模型初始化时指定dtype。
  • 清理CUDA缓存:训练前执行torch.cuda.empty_cache()释放无用显存,避免显存碎片导致异常。
  • 验证CUDA环境:执行以下代码确认PyTorch能正确识别CUDA:
import torch
print(torch.cuda.is_available())
print(torch.cuda.get_device_name(0))

若返回True和GPU名称,说明CUDA环境正常;否则需重新安装匹配版本的PyTorch和CUDA。

  • 调整PyTorch Lightning版本:1.7.7版本可能存在兼容问题,可尝试升级到1.8.x或降级到1.6.x版本,看是否解决问题。

内容的提问来源于stack exchange,提问作者Ysk196

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 18:10:35