You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

RTX 2080Ti本地微调Llama遇RuntimeError:is_sm80||is_sm90验证失败

解决RTX 2080Ti(sm75)微调Llama时的Expected is_sm80 || is_sm90错误

错误原因

这个错误是因为训练代码中调用了仅支持NVIDIA Ampere(sm80)及以上架构的CUDA算子(比如FlashAttention v2、PyTorch 2.1+新增的部分优化算子),而RTX 2080Ti属于Turing架构(sm75),不支持这些算子。

可行解决方案

1. 禁用FlashAttention

大部分LLM训练脚本默认启用FlashAttention加速,这是触发错误的核心原因,你可以通过以下方式关闭:

  • 在训练命令中添加参数:--disable_flash_attn
  • 在代码中手动配置:
    model.config.use_flash_attention_2 = False
    
  • 使用SFTTrainer时,在初始化参数中设置:
    from transformers import TrainingArguments
    args = TrainingArguments(..., disable_flash_attn=True)
    

2. 降级PyTorch到2.0.x版本

PyTorch 2.1及以上版本对部分算子强制要求sm80+架构,降级到2.0.x版本可恢复对sm75的支持。根据你的CUDA版本选择安装命令,比如适配CUDA 11.8的命令:

pip install torch==2.0.1 torchvision==0.15.2 torchaudio==2.0.2 --index-url https://download.pytorch.org/whl/cu118

注意:RTX 2080Ti最高支持CUDA 11.8,不要使用CUDA 12.x版本。

3. 适配依赖库版本

使用对老GPU更友好的transformers、trl库版本:

pip install transformers==4.34.1 trl==0.7.4 accelerate==0.23.0

4. 替换高架构依赖算子

如果代码中直接使用了torch.nn.functional.scaled_dot_product_attention,可以替换为传统注意力实现,或者添加架构判断逻辑:

import math
import torch

def custom_attention(query, key, value):
    # 检测GPU架构,低于sm80时使用传统注意力
    if torch.cuda.get_device_capability()[0] < 8:
        attn_weights = torch.matmul(query, key.transpose(-2, -1)) / math.sqrt(query.size(-1))
        attn_weights = torch.softmax(attn_weights, dim=-1)
        return torch.matmul(attn_weights, value)
    else:
        return torch.nn.functional.scaled_dot_product_attention(query, key, value)

显卡兼容性说明

RTX 2080Ti完全支持Llama的微调任务,只是无法使用仅Ampere及以上架构支持的新优化算子。调整上述配置后,训练可以正常运行,只是速度会比Ampere架构显卡慢一些。

内容的提问来源于stack exchange,提问作者Tr33Bug

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 16:52:33