You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Kosmos-2转ONNX后离线推理输出异常问题求助

Kosmos-2转ONNX后推理结果异常的解决方法

我想把HuggingFace的Kosmos-2模型转成ONNX格式,用于Flutter应用的离线推理,但转换后推理结果完全不符合预期:输出的last_hidden_state形状与预期不符,且结果为浮点数而非0-64000范围内的整数。以下是我的代码、输出及解决方案:


输入输出信息查看代码

inputs = sess.get_inputs()
outputs = sess.get_outputs()

for i, input_info in enumerate(inputs):
    print(f"Input {i}: {input_info.name} shape: {input_info.shape}")

#for i in range(len(sess.get_outputs())):
#output_name = sess.get_outputs()[i].name
#print("Output name: ", output_name)
#output names: only last_hidden_state
for i, output_info in enumerate(inputs):
    print(f"Output:", {output_info})
output_name = sess.get_outputs()[0]
print("Output 1:", output_name)
input_name = sess.get_inputs()[0]

对应输出

Input 0: pixel_values shape: ['batch_size', 'num_channels', 'height', 'width']
Input 1: input_ids shape: ['batch_size', 'sequence_length']
Input 2: attention_mask shape: ['batch_size', 'sequence_length']
Input 3: image_embeds_position_mask shape: ['batch_size', 'sequence_length']
Output: {<onnxruntime.capi.onnxruntime_pybind11_state.NodeArg object at 0x000002A0462509B0>}
Output: {<onnxruntime.capi.onnxruntime_pybind11_state.NodeArg object at 0x000002A0120D1830>}
Output: {<onnxruntime.capi.onnxruntime_pybind11_state.NodeArg object at 0x000002A046F3E230>}
Output: {<onnxruntime.capi.onnxruntime_pybind11_state.NodeArg object at 0x000002A046F3DAB0>}
Output 1: NodeArg(name='last_hidden_state', type='tensor(float)', shape=['batch_size', 'sequence_length', 'hidden_size'])

推理代码

tokenInputs = tokenizer(text="A picture of", images=image, return_tensors="pt")
processor = AutoProcessor.from_pretrained("microsoft/kosmos-2-patch14-224")
print(tokenInputs['pixel_values'].shape)
#print("Token inputs:", tokenInputs)
print(tokenInputs['input_ids'].to(dtype=torch.int64).shape)
example = torch.zeros(1, 75)
#out_mat = sess.run(output_names=['last_hidden_state'], input_feed={'pixel_values': np.array(tokenInputs['pixel_values']), 'input_ids': np.array(tokenInputs['input_ids'].to(dtype=torch.int64)), 'attention_mask': np.array(tokenInputs['attention_mask'].to(dtype=torch.int64)), 'image_embeds_position_mask': np.array(tokenInputs['image_embeds_position_mask'].to(dtype=torch.int64))})
out_mat = sess.run(output_names=['last_hidden_state'], input_feed={'pixel_values':np.array(tokenInputs['pixel_values']), 'input_ids':np.array(example.to(dtype=torch.int64)), 'attention_mask': np.array(example.to(dtype=torch.int64)), 'image_embeds_position_mask': np.array(example.to(dtype=torch.int64))})
out_mat = torch.tensor(out_mat)
print(out_mat.shape)
print(out_mat)

对应输出

Special tokens have been added in the vocabulary, make sure the associated word embeddings are fine-tuned or trained.
torch.Size([1, 1, 75, 2048])
tensor([[[[ 0.0831, -0.1343, -0.3462, ..., 0.5155, -0.2754, 0.0927],
[ 0.0808, -0.1369, -0.3498, ..., 0.5168, -0.2769, 0.0934],
[ 0.0780, -0.1400, -0.3541, ..., 0.5188, -0.2782, 0.0938],
...,
[ 0.0746, -0.1454, -0.3643, ..., 0.5329, -0.2984, 0.0831],
[ 0.0780, -0.1452, -0.3672, ..., 0.5332, -0.2987, 0.0835],
[ 0.0834, -0.1424, -0.3674, ..., 0.5337, -0.2981, 0.0835]]]])

问题分析与解决方法

1. 输出形状异常问题

当前输出形状为[1,1,75,2048],多了一层维度,原因是onnxruntime的sess.run()返回的是包含输出张量的列表,直接转成torch.tensor会保留列表的外层维度。

解决方法:取列表第一个元素再转成tensor:

out_mat = torch.tensor(out_mat[0])

处理后形状会变为[1,75,2048],符合[batch_size, sequence_length, hidden_size]的预期。

2. 输出为浮点数而非整数问题

last_hidden_state是模型的隐藏层特征向量,本身就是浮点数;你需要的0-64000范围内的整数是模型经过**语言模型头(LM Head)**处理后,对logits取argmax得到的token ID。问题出在导出ONNX时只导出了模型编码器部分,未包含LM Head。

步骤1:导出包含LM Head的完整模型

确保导出的是带LM Head的完整模型,示例代码:

from transformers import AutoModelForVisionAndLanguageGeneration
import torch

model = AutoModelForVisionAndLanguageGeneration.from_pretrained("microsoft/kosmos-2-patch14-224")
processor = AutoProcessor.from_pretrained("microsoft/kosmos-2-patch14-224")

# 构造示例输入
text = "A picture of"
image = ... # 你的输入图像
inputs = processor(text=text, images=image, return_tensors="pt")

# 导出完整模型,指定输出为logits
torch.onnx.export(
    model,
    tuple(inputs.values()),
    "kosmos2_full.onnx",
    opset_version=17,
    input_names=['pixel_values', 'input_ids', 'attention_mask', 'image_embeds_position_mask'],
    output_names=['logits'],
    dynamic_axes={
        'input_ids': {0: 'batch_size', 1: 'sequence_length'},
        'attention_mask': {0: 'batch_size', 1: 'sequence_length'},
        'image_embeds_position_mask': {0: 'batch_size', 1: 'sequence_length'},
        'pixel_values': {0: 'batch_size'},
        'logits': {0: 'batch_size', 1: 'sequence_length'}
    }
)

步骤2:推理时从logits获取token ID

导出后,推理得到logits后对最后一维取argmax,即可得到目标整数:

import onnxruntime as ort
import numpy as np
import torch

sess = ort.InferenceSession("kosmos2_full.onnx")
# 使用processor处理后的真实输入,而非全零张量
tokenInputs = processor(text="A picture of", images=image, return_tensors="pt")
input_feed = {
    'pixel_values': np.array(tokenInputs['pixel_values']),
    'input_ids': np.array(tokenInputs['input_ids'].to(torch.int64)),
    'attention_mask': np.array(tokenInputs['attention_mask'].to(torch.int64)),
    'image_embeds_position_mask': np.array(tokenInputs['image_embeds_position_mask'].to(torch.int64))
}
logits = sess.run(output_names=['logits'], input_feed=input_feed)[0]
# 取每个位置的最大概率token ID
token_ids = np.argmax(logits, axis=-1)
print(token_ids) # 输出为0-64000范围内的整数

3. 额外注意事项

  • 推理时不能用全零张量作为input_ids等输入,必须使用processor处理后的真实输入,否则输出无意义。
  • 导出ONNX时需指定dynamic_axes,避免输入尺寸固定导致后续推理报错。
  • 确保ONNX Runtime与PyTorch版本兼容,建议使用较新版本(如ORT 1.15+、PyTorch 2.0+)。

内容的提问来源于stack exchange,提问作者user23294494

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 15:35:18