LLaVA-1.6-Mistral-7B模型推理返回<UNK>异常排查求助
LLaVA-1.6-Mistral-7B模型推理返回异常排查求助
我最近在尝试用llava-hf/llava-v1.6-mistral-7b-hf模型提取图片中的信息和文本,但不管怎么调整,模型输出结果全是一堆<unk>占位符,完全无法正常识别图片内容,输出示例如下:
[/INST]<unk><unk><unk> <unk>
下面是我完整的运行代码,麻烦各位帮我看看哪里出问题了:
import torch from PIL import Image import requests import numpy as np # 设备配置 if torch.cuda.is_available(): device = torch.device("cuda:0") print("running on the GPU") else: device = torch.device("cpu") print("running on the CPU") processor = LlavaNextProcessor.from_pretrained(r"D:\DL_ML_Projects\huggingface\hub\models--llava-hf--llava-v1.6-mistral-7b-hf\snapshots\ed45ac3a134a109ea7148e5ba3639d8972cad49c") model = LlavaNextForConditionalGeneration.from_pretrained(r"D:\DL_ML_Projects\huggingface\hub\models--llava-hf--llava-v1.6-mistral-7b-hf\snapshots\ed45ac3a134a109ea7148e5ba3639d8972cad49c", torch_dtype=torch.float16, low_cpu_mem_usage=True) model = model.to("cuda:0") # 加载图片 image = Image.open(r"page_354.jpg").convert("RGB") # 构建对话模板 conversation = [ { "role": "user", "content": [ {"type": "text", "text": "get the information shown in the image?"}, {"type": "image"}, ], } ] prompt = processor.apply_chat_template(conversation, add_generation_prompt=True) # 准备输入并推理 inputs = processor(text=prompt, images=image, return_tensors="pt").to("cuda:0") output = model.generate(**inputs, pad_token_id=model.config.eos_token_id, max_new_tokens=512) # 打印输出 print(processor.decode(output[0], skip_special_tokens=False))
我已经确认模型文件是从Hugging Face Hub下载的完整快照,CUDA环境也正常能识别到GPU,但就是解决不了
备注:内容来源于stack exchange,提问作者Mohamed Hassan
相关产品推荐
相关产品推荐

