在Kaggle Notebook运行Hugging Face Llama3模型时遇RuntimeError
问题:Kaggle Notebook调用Llama3 pipeline触发cutlassF内核错误
在Kaggle Notebook中使用Hugging Face的Llama3模型,调用pipeline模块时出现如下核心报错:
RuntimeError Traceback (most recent call last) Cell In[19], line 17, in Llama_Chat(system_role, user_msg) 12 def Llama_Chat(system_role,user_msg): 13 messages = [ 14 {"role": "system", "content": system_role}, 15 {"role": "user", "content": user_msg}, 16 ] ---> 17 outputs = pipeline( 18 messages, 19 max_new_tokens=256, 20 temperature = 0.1 21 22 ) 24 reply=outputs[0]["generated_text"][-1]["content"] 25 return reply File /opt/conda/lib/python3.10/site-packages/accelerate/hooks.py:169, in add_hook_to_module.<locals>.new_forward(module, *args, **kwargs) 167 output = module._old_forward(*args, **kwargs) 168 else: --> 169 output = module._old_forward(*args, **kwargs) 170 return module._hf_hook.post_forward(module, output) File /opt/conda/lib/python3.10/site-packages/transformers/models/llama/modeling_llama.py:603, in LlamaSdpaAttention.forward(self, hidden_states, attention_mask, position_ids, past_key_value, output_attentions, use_cache, cache_position, position_embeddings, **kwargs) 599 # We dispatch to SDPA's Flash Attention or Efficient kernels via this `is_causal` if statement instead of an inline conditional assignment 600 # in SDPA to support both torch.compile's dynamic shapes and full graph options. An inline conditional prevents dynamic shapes from compiling. 601 is_causal = True if causal_mask is None and q_len > 1 else False --> 603 attn_output = torch.nn.functional.scaled_dot_product_attention( 604 query_states, 605 key_states, 606 value_states, 607 attn_mask=causal_mask, 608 dropout_p=self.attention_dropout if self.training else 0.0, 609 is_causal=is_causal, 610 ) 612 attn_output = attn_output.transpose(1, 2).contiguous() 613 attn_output = attn_output.view(bsz, q_len, -1) RuntimeError: cutlassF: no kernel found to launch!
已确认CUDA、PyTorch版本无异常,但常规AI工具仅提示版本不兼容,无法推进解决。
解决方案
禁用Flash Attention 2:这是最直接的修复方式,报错源于Scaled Dot-Product Attention(SDPA)尝试调用不兼容的Cutlass内核。可以在初始化pipeline或模型时强制使用eager模式的注意力实现:
# 方式1:初始化pipeline时指定 from transformers import pipeline pipe = pipeline( "conversational", model="meta-llama/Meta-Llama-3-8B-Instruct", use_flash_attention_2=False, device_map="auto" ) # 方式2:加载模型时指定注意力实现 from transformers import AutoModelForCausalLM, AutoTokenizer model = AutoModelForCausalLM.from_pretrained( "meta-llama/Meta-Llama-3-8B-Instruct", attn_implementation="eager", device_map="auto" ) tokenizer = AutoTokenizer.from_pretrained("meta-llama/Meta-Llama-3-8B-Instruct")锁定兼容版本组合:即使你认为版本无异常,部分版本组合可能存在隐性兼容性问题。尝试安装经过验证的版本:
!pip install torch==2.1.2 transformers==4.38.2 accelerate==0.27.2检查GPU实例兼容性:Kaggle提供的部分GPU(如T4)对Flash Attention的支持有限,切换到A100等更高规格的GPU实例测试是否解决问题。
调整推理参数:尝试减小
max_new_tokens值,或者确保输入张量的形状符合Cutlass内核的要求,避免因非常规形状导致找不到匹配的内核。
内容的提问来源于stack exchange,提问作者shivam
相关产品推荐
相关产品推荐

