Kosmos-2转ONNX后离线推理输出异常问题求助
Kosmos-2转ONNX后推理结果异常的解决方法
我想把HuggingFace的Kosmos-2模型转成ONNX格式,用于Flutter应用的离线推理,但转换后推理结果完全不符合预期:输出的last_hidden_state形状与预期不符,且结果为浮点数而非0-64000范围内的整数。以下是我的代码、输出及解决方案:
输入输出信息查看代码
inputs = sess.get_inputs() outputs = sess.get_outputs() for i, input_info in enumerate(inputs): print(f"Input {i}: {input_info.name} shape: {input_info.shape}") #for i in range(len(sess.get_outputs())): #output_name = sess.get_outputs()[i].name #print("Output name: ", output_name) #output names: only last_hidden_state for i, output_info in enumerate(inputs): print(f"Output:", {output_info}) output_name = sess.get_outputs()[0] print("Output 1:", output_name) input_name = sess.get_inputs()[0]
对应输出
Input 0: pixel_values shape: ['batch_size', 'num_channels', 'height', 'width'] Input 1: input_ids shape: ['batch_size', 'sequence_length'] Input 2: attention_mask shape: ['batch_size', 'sequence_length'] Input 3: image_embeds_position_mask shape: ['batch_size', 'sequence_length'] Output: {<onnxruntime.capi.onnxruntime_pybind11_state.NodeArg object at 0x000002A0462509B0>} Output: {<onnxruntime.capi.onnxruntime_pybind11_state.NodeArg object at 0x000002A0120D1830>} Output: {<onnxruntime.capi.onnxruntime_pybind11_state.NodeArg object at 0x000002A046F3E230>} Output: {<onnxruntime.capi.onnxruntime_pybind11_state.NodeArg object at 0x000002A046F3DAB0>} Output 1: NodeArg(name='last_hidden_state', type='tensor(float)', shape=['batch_size', 'sequence_length', 'hidden_size'])
推理代码
tokenInputs = tokenizer(text="A picture of", images=image, return_tensors="pt") processor = AutoProcessor.from_pretrained("microsoft/kosmos-2-patch14-224") print(tokenInputs['pixel_values'].shape) #print("Token inputs:", tokenInputs) print(tokenInputs['input_ids'].to(dtype=torch.int64).shape) example = torch.zeros(1, 75) #out_mat = sess.run(output_names=['last_hidden_state'], input_feed={'pixel_values': np.array(tokenInputs['pixel_values']), 'input_ids': np.array(tokenInputs['input_ids'].to(dtype=torch.int64)), 'attention_mask': np.array(tokenInputs['attention_mask'].to(dtype=torch.int64)), 'image_embeds_position_mask': np.array(tokenInputs['image_embeds_position_mask'].to(dtype=torch.int64))}) out_mat = sess.run(output_names=['last_hidden_state'], input_feed={'pixel_values':np.array(tokenInputs['pixel_values']), 'input_ids':np.array(example.to(dtype=torch.int64)), 'attention_mask': np.array(example.to(dtype=torch.int64)), 'image_embeds_position_mask': np.array(example.to(dtype=torch.int64))}) out_mat = torch.tensor(out_mat) print(out_mat.shape) print(out_mat)
对应输出
Special tokens have been added in the vocabulary, make sure the associated word embeddings are fine-tuned or trained. torch.Size([1, 1, 75, 2048]) tensor([[[[ 0.0831, -0.1343, -0.3462, ..., 0.5155, -0.2754, 0.0927], [ 0.0808, -0.1369, -0.3498, ..., 0.5168, -0.2769, 0.0934], [ 0.0780, -0.1400, -0.3541, ..., 0.5188, -0.2782, 0.0938], ..., [ 0.0746, -0.1454, -0.3643, ..., 0.5329, -0.2984, 0.0831], [ 0.0780, -0.1452, -0.3672, ..., 0.5332, -0.2987, 0.0835], [ 0.0834, -0.1424, -0.3674, ..., 0.5337, -0.2981, 0.0835]]]])
问题分析与解决方法
1. 输出形状异常问题
当前输出形状为[1,1,75,2048],多了一层维度,原因是onnxruntime的sess.run()返回的是包含输出张量的列表,直接转成torch.tensor会保留列表的外层维度。
解决方法:取列表第一个元素再转成tensor:
out_mat = torch.tensor(out_mat[0])
处理后形状会变为[1,75,2048],符合[batch_size, sequence_length, hidden_size]的预期。
2. 输出为浮点数而非整数问题
last_hidden_state是模型的隐藏层特征向量,本身就是浮点数;你需要的0-64000范围内的整数是模型经过**语言模型头(LM Head)**处理后,对logits取argmax得到的token ID。问题出在导出ONNX时只导出了模型编码器部分,未包含LM Head。
步骤1:导出包含LM Head的完整模型
确保导出的是带LM Head的完整模型,示例代码:
from transformers import AutoModelForVisionAndLanguageGeneration import torch model = AutoModelForVisionAndLanguageGeneration.from_pretrained("microsoft/kosmos-2-patch14-224") processor = AutoProcessor.from_pretrained("microsoft/kosmos-2-patch14-224") # 构造示例输入 text = "A picture of" image = ... # 你的输入图像 inputs = processor(text=text, images=image, return_tensors="pt") # 导出完整模型,指定输出为logits torch.onnx.export( model, tuple(inputs.values()), "kosmos2_full.onnx", opset_version=17, input_names=['pixel_values', 'input_ids', 'attention_mask', 'image_embeds_position_mask'], output_names=['logits'], dynamic_axes={ 'input_ids': {0: 'batch_size', 1: 'sequence_length'}, 'attention_mask': {0: 'batch_size', 1: 'sequence_length'}, 'image_embeds_position_mask': {0: 'batch_size', 1: 'sequence_length'}, 'pixel_values': {0: 'batch_size'}, 'logits': {0: 'batch_size', 1: 'sequence_length'} } )
步骤2:推理时从logits获取token ID
导出后,推理得到logits后对最后一维取argmax,即可得到目标整数:
import onnxruntime as ort import numpy as np import torch sess = ort.InferenceSession("kosmos2_full.onnx") # 使用processor处理后的真实输入,而非全零张量 tokenInputs = processor(text="A picture of", images=image, return_tensors="pt") input_feed = { 'pixel_values': np.array(tokenInputs['pixel_values']), 'input_ids': np.array(tokenInputs['input_ids'].to(torch.int64)), 'attention_mask': np.array(tokenInputs['attention_mask'].to(torch.int64)), 'image_embeds_position_mask': np.array(tokenInputs['image_embeds_position_mask'].to(torch.int64)) } logits = sess.run(output_names=['logits'], input_feed=input_feed)[0] # 取每个位置的最大概率token ID token_ids = np.argmax(logits, axis=-1) print(token_ids) # 输出为0-64000范围内的整数
3. 额外注意事项
- 推理时不能用全零张量作为
input_ids等输入,必须使用processor处理后的真实输入,否则输出无意义。 - 导出ONNX时需指定
dynamic_axes,避免输入尺寸固定导致后续推理报错。 - 确保ONNX Runtime与PyTorch版本兼容,建议使用较新版本(如ORT 1.15+、PyTorch 2.0+)。
内容的提问来源于stack exchange,提问作者user23294494
相关产品推荐
相关产品推荐

