You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

jinaai/jina-embeddings-v3模型无法输出attention问题求助

问题分析与解决

问题原因

jinaai/jina-embeddings-v3的自定义模型代码存在参数处理逻辑缺陷:

  • 未安装flash_attn时,模型会切换到PyTorch原生注意力实现,但原生实现的逻辑中并未正确响应output_attentions=True参数,导致即使设置该参数,也无法返回注意力权重。
  • 警告信息矛盾是因为模型代码在检测flash_attn状态时,错误输出了针对Flash Attention的不支持提示,即便此时实际使用的是原生注意力。

解决方案

方案1:手动修复模型注意力层逻辑

加载模型后,修改注意力层的forward方法,让原生注意力实现返回注意力权重:

from transformers import AutoModel, AutoTokenizer

model_id = "jinaai/jina-embeddings-v3"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModel.from_pretrained(model_id, trust_remote_code=True)

# 重写注意力层的forward方法,支持返回注意力权重
def patched_attn_forward(self, hidden_states, attention_mask=None, output_attentions=False):
    attn_output, attn_weights = self.attn(hidden_states, attention_mask=attention_mask, return_attentions=output_attentions)
    if output_attentions:
        return attn_output, attn_weights
    return attn_output

# 遍历替换所有层的注意力forward方法
for layer in model.layers:
    layer.attention.forward = patched_attn_forward.__get__(layer.attention, type(layer.attention))

# 重新执行推理
inputs = tokenizer([
    "The weather is lovely today.",
    "It's so sunny outside!",
    "He drove to the stadium."
], return_tensors="pt", padding=True, truncation=True)

outputs = model(**inputs, output_attentions=True)
attentions = outputs.attentions
print(attentions)  # 此时可正常获取注意力权重

方案2:强制禁用Flash Attention

加载模型时,通过参数强制禁用Flash Attention,确保使用原生实现并尝试触发注意力返回:

model = AutoModel.from_pretrained(
    model_id,
    trust_remote_code=True,
    use_flash_attention=False  # 强制使用原生注意力
)

注:若模型代码未暴露use_flash_attention参数,需手动修改模型配置文件中的对应设置。

方案3:等待官方修复

该问题属于模型自定义代码的bug,可在模型的Hugging Face页面提交issue,等待官方更新代码以正确处理output_attentions参数。

内容的提问来源于stack exchange,提问作者Yash Mali

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 13:54:58