修复TORCH_USE_CUDA_DSA运行时错误时PyTorch重装失败
问题与解决方案
问题描述
- 运行一款接收用户输入生成字符集的LLM程序时,触发CUDA相关错误提示:
For debugging consider passing CUDA_LAUNCH_BLOCKING=1. Compile with 'TORCH_USE_CUDA_DSA' to enable device-side assertions. - 尝试通过以下命令重装PyTorch nightly版本失败:
报错信息:pip install --pre torch torchvision torchaudio --force-reinstall --index-url https://download.pytorch.org/whl/nightly/cu118ERROR: Could not find a version that satisfies the requirement torchvision (from versions: none) ERROR: No matching distribution found for torchvision - 单独重装不含torchvision、torchaudio的Torch后,原CUDA错误仍未消失。
解决步骤
1. 排查CUDA环境匹配问题
- 用
nvcc --version查看本地CUDA版本,确认是否为11.8;若版本不匹配,更换对应PyTorch源:- 比如CUDA 12.1对应源
https://download.pytorch.org/whl/nightly/cu121 - CPU环境对应源
https://download.pytorch.org/whl/nightly/cpu
- 比如CUDA 12.1对应源
2. 修复torchvision安装失败问题
- 先查询对应CUDA版本的PyTorch nightly可用版本:
拿到版本号(如pip index versions torch --index-url https://download.pytorch.org/whl/nightly/cu1182.3.0.dev20240520+cu118)后,指定完整版本号重装:pip install --pre torch==2.3.0.dev20240520+cu118 torchvision==0.18.0.dev20240520+cu118 torchaudio==2.3.0.dev20240520+cu118 --force-reinstall --index-url https://download.pytorch.org/whl/nightly/cu118 - 若仍失败,先清理现有组件再重装稳定版:
pip uninstall -y torch torchvision torchaudio pip cache purge pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu118
3. 定位并解决原CUDA运行错误
- 开启调试模式获取详细错误:
重新运行程序,根据输出的错误栈定位具体问题。# Linux/macOS export CUDA_LAUNCH_BLOCKING=1 export TORCH_USE_CUDA_DSA=1 # Windows cmd set CUDA_LAUNCH_BLOCKING=1 set TORCH_USE_CUDA_DSA=1 # Windows PowerShell $env:CUDA_LAUNCH_BLOCKING=1 $env:TORCH_USE_CUDA_DSA=1 - 常见问题修复:
- 内存不足:减小batch size、启用模型4bit/8bit量化、调用
torch.cuda.empty_cache()清理GPU缓存 - 张量形状不兼容:检查用户输入的字符集张量维度是否匹配模型要求
- 驱动版本过低:更新显卡驱动到对应CUDA版本要求的最低版本
- 内存不足:减小batch size、启用模型4bit/8bit量化、调用
内容的提问来源于stack exchange,提问作者Leon Tipton
相关产品推荐
相关产品推荐

