You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

LLaVA-1.6-Mistral-7B模型推理返回<UNK>异常排查求助

LLaVA-1.6-Mistral-7B模型推理返回异常排查求助

我最近在尝试用llava-hf/llava-v1.6-mistral-7b-hf模型提取图片中的信息和文本,但不管怎么调整,模型输出结果全是一堆<unk>占位符,完全无法正常识别图片内容,输出示例如下:

[/INST]<unk><unk><unk> <unk>

下面是我完整的运行代码,麻烦各位帮我看看哪里出问题了:

import torch
from PIL import Image
import requests
import numpy as np

# 设备配置
if torch.cuda.is_available():
    device = torch.device("cuda:0")
    print("running on the GPU")
else:
    device = torch.device("cpu")
    print("running on the CPU")

processor = LlavaNextProcessor.from_pretrained(r"D:\DL_ML_Projects\huggingface\hub\models--llava-hf--llava-v1.6-mistral-7b-hf\snapshots\ed45ac3a134a109ea7148e5ba3639d8972cad49c")
model = LlavaNextForConditionalGeneration.from_pretrained(r"D:\DL_ML_Projects\huggingface\hub\models--llava-hf--llava-v1.6-mistral-7b-hf\snapshots\ed45ac3a134a109ea7148e5ba3639d8972cad49c", 
                                                          torch_dtype=torch.float16, 
                                                          low_cpu_mem_usage=True)
model = model.to("cuda:0")

# 加载图片
image = Image.open(r"page_354.jpg").convert("RGB")

# 构建对话模板
conversation = [ 
    { 
        "role": "user", 
        "content": [ 
            {"type": "text", "text": "get the information shown in the image?"}, 
            {"type": "image"}, 
        ], 
    }  
] 

prompt = processor.apply_chat_template(conversation, add_generation_prompt=True)

# 准备输入并推理
inputs = processor(text=prompt, images=image, return_tensors="pt").to("cuda:0")
output = model.generate(**inputs, pad_token_id=model.config.eos_token_id, max_new_tokens=512)

# 打印输出
print(processor.decode(output[0], skip_special_tokens=False))

我已经确认模型文件是从Hugging Face Hub下载的完整快照,CUDA环境也正常能识别到GPU,但就是解决不了的问题,有没有大佬能指点一下可能的原因和解决办法?

备注:内容来源于stack exchange,提问作者Mohamed Hassan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 18:19:28