You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Kaggle Notebook运行Hugging Face Llama3模型时遇RuntimeError

问题:Kaggle Notebook调用Llama3 pipeline触发cutlassF内核错误

在Kaggle Notebook中使用Hugging Face的Llama3模型,调用pipeline模块时出现如下核心报错:

RuntimeError                              Traceback (most recent call last)
Cell In[19], line 17, in Llama_Chat(system_role, user_msg)
     12 def Llama_Chat(system_role,user_msg):
     13   messages = [
     14     {"role": "system", "content": system_role},
     15     {"role": "user", "content": user_msg},
     16   ]
---> 17   outputs = pipeline(
     18       messages,
     19       max_new_tokens=256,
     20       temperature = 0.1
     21 
     22   )
     24   reply=outputs[0]["generated_text"][-1]["content"]
     25   return reply

File /opt/conda/lib/python3.10/site-packages/accelerate/hooks.py:169, in add_hook_to_module.<locals>.new_forward(module, *args, **kwargs)
    167         output = module._old_forward(*args, **kwargs)
    168 else:
--> 169     output = module._old_forward(*args, **kwargs)
    170 return module._hf_hook.post_forward(module, output)

File /opt/conda/lib/python3.10/site-packages/transformers/models/llama/modeling_llama.py:603, in LlamaSdpaAttention.forward(self, hidden_states, attention_mask, position_ids, past_key_value, output_attentions, use_cache, cache_position, position_embeddings, **kwargs)
    599 # We dispatch to SDPA's Flash Attention or Efficient kernels via this `is_causal` if statement instead of an inline conditional assignment
    600 # in SDPA to support both torch.compile's dynamic shapes and full graph options. An inline conditional prevents dynamic shapes from compiling.
    601 is_causal = True if causal_mask is None and q_len > 1 else False
--> 603 attn_output = torch.nn.functional.scaled_dot_product_attention(
    604     query_states,
    605     key_states,
    606     value_states,
    607     attn_mask=causal_mask,
    608     dropout_p=self.attention_dropout if self.training else 0.0,
    609     is_causal=is_causal,
    610 )
    612 attn_output = attn_output.transpose(1, 2).contiguous()
    613 attn_output = attn_output.view(bsz, q_len, -1)

RuntimeError: cutlassF: no kernel found to launch!

已确认CUDA、PyTorch版本无异常,但常规AI工具仅提示版本不兼容,无法推进解决。


解决方案
  • 禁用Flash Attention 2:这是最直接的修复方式,报错源于Scaled Dot-Product Attention(SDPA)尝试调用不兼容的Cutlass内核。可以在初始化pipeline或模型时强制使用eager模式的注意力实现:

    # 方式1:初始化pipeline时指定
    from transformers import pipeline
    pipe = pipeline(
        "conversational",
        model="meta-llama/Meta-Llama-3-8B-Instruct",
        use_flash_attention_2=False,
        device_map="auto"
    )
    
    # 方式2:加载模型时指定注意力实现
    from transformers import AutoModelForCausalLM, AutoTokenizer
    model = AutoModelForCausalLM.from_pretrained(
        "meta-llama/Meta-Llama-3-8B-Instruct",
        attn_implementation="eager",
        device_map="auto"
    )
    tokenizer = AutoTokenizer.from_pretrained("meta-llama/Meta-Llama-3-8B-Instruct")
    
  • 锁定兼容版本组合:即使你认为版本无异常,部分版本组合可能存在隐性兼容性问题。尝试安装经过验证的版本:

    !pip install torch==2.1.2 transformers==4.38.2 accelerate==0.27.2
    
  • 检查GPU实例兼容性:Kaggle提供的部分GPU(如T4)对Flash Attention的支持有限,切换到A100等更高规格的GPU实例测试是否解决问题。

  • 调整推理参数:尝试减小max_new_tokens值,或者确保输入张量的形状符合Cutlass内核的要求,避免因非常规形状导致找不到匹配的内核。


内容的提问来源于stack exchange,提问作者shivam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 13:34:57