RTX 2080Ti本地微调Llama遇RuntimeError:is_sm80||is_sm90验证失败
解决RTX 2080Ti(sm75)微调Llama时的
Expected is_sm80 || is_sm90错误 错误原因
这个错误是因为训练代码中调用了仅支持NVIDIA Ampere(sm80)及以上架构的CUDA算子(比如FlashAttention v2、PyTorch 2.1+新增的部分优化算子),而RTX 2080Ti属于Turing架构(sm75),不支持这些算子。
可行解决方案
1. 禁用FlashAttention
大部分LLM训练脚本默认启用FlashAttention加速,这是触发错误的核心原因,你可以通过以下方式关闭:
- 在训练命令中添加参数:
--disable_flash_attn - 在代码中手动配置:
model.config.use_flash_attention_2 = False - 使用SFTTrainer时,在初始化参数中设置:
from transformers import TrainingArguments args = TrainingArguments(..., disable_flash_attn=True)
2. 降级PyTorch到2.0.x版本
PyTorch 2.1及以上版本对部分算子强制要求sm80+架构,降级到2.0.x版本可恢复对sm75的支持。根据你的CUDA版本选择安装命令,比如适配CUDA 11.8的命令:
pip install torch==2.0.1 torchvision==0.15.2 torchaudio==2.0.2 --index-url https://download.pytorch.org/whl/cu118
注意:RTX 2080Ti最高支持CUDA 11.8,不要使用CUDA 12.x版本。
3. 适配依赖库版本
使用对老GPU更友好的transformers、trl库版本:
pip install transformers==4.34.1 trl==0.7.4 accelerate==0.23.0
4. 替换高架构依赖算子
如果代码中直接使用了torch.nn.functional.scaled_dot_product_attention,可以替换为传统注意力实现,或者添加架构判断逻辑:
import math import torch def custom_attention(query, key, value): # 检测GPU架构,低于sm80时使用传统注意力 if torch.cuda.get_device_capability()[0] < 8: attn_weights = torch.matmul(query, key.transpose(-2, -1)) / math.sqrt(query.size(-1)) attn_weights = torch.softmax(attn_weights, dim=-1) return torch.matmul(attn_weights, value) else: return torch.nn.functional.scaled_dot_product_attention(query, key, value)
显卡兼容性说明
RTX 2080Ti完全支持Llama的微调任务,只是无法使用仅Ampere及以上架构支持的新优化算子。调整上述配置后,训练可以正常运行,只是速度会比Ampere架构显卡慢一些。
内容的提问来源于stack exchange,提问作者Tr33Bug
相关产品推荐
相关产品推荐

