同一LLaVA模型在不同Transformers版本中的分词与生成差异求助
问题描述
我下载了一个基于LLaVA的自定义旧模型,该模型原本适配Transformers 4.31.0版本。由于需要搭配使用适配Transformers 4.53.1版本的Qwen模型,我升级了Transformers版本,此后该LLaVA模型无法正常运行。经排查发现,即使模型及配置完全相同,同一prompt在不同Transformers版本下的分词与生成行为仍存在差异。
差异示例
测试Prompt
"A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <im_start><image><im_end>\nPlease locate the tape in this image. ASSISTANT:"
分词结果差异
- Transformers 4.31.0分词结果:

- Transformers 4.53.1分词结果:

模型输出差异
- Transformers 4.31.0输出:

- Transformers 4.53.1输出:

疑问
我想了解为何会出现这种模型行为差异,以及是否可以让该模型恢复在Transformers 4.31.0版本下的运行行为。
示例代码
from transformers import AutoTokenizer, CLIPImageProcessor import torch from PIL import Image kwargs = {} kwargs['torch_dtype'] = torch.bfloat16 kwargs['device_map'] = 'cuda' kwargs['is_eval'] = True vsm_tokenizer = AutoTokenizer.from_pretrained( 'craigwu/seal_vsm_7b', cache_dir=None, model_max_length=512, padding_side="right", use_fast=False ) vsm_tokenizer.pad_token = vsm_tokenizer.unk_token loc_token_idx = vsm_tokenizer("[LOC]", add_special_tokens=False).input_ids[0] vsm_model = VSMForCausalLM.from_pretrained( 'craigwu/seal_vsm_7b', low_cpu_mem_usage=True, vision_tower='openai/clip-vit-large-patch14', loc_token_idx=loc_token_idx, **kwargs ) vsm_model.get_model().initialize_vision_modules(vsm_model.get_model().config) vsm_model.eval() clip_image_processor = CLIPImageProcessor.from_pretrained('openai/clip-vit-large-patch14') def tokenizer_image_token( prompt, tokenizer, image_token_index=IMAGE_TOKEN_INDEX, return_tensors=None ): prompt_chunks = [tokenizer(chunk).input_ids for chunk in prompt.split("<image>")] def insert_separator(X, sep): return [ele for sublist in zip(X, [sep] * len(X)) for ele in sublist][:-1] input_ids = [] offset = 0 if ( len(prompt_chunks) > 0 and len(prompt_chunks[0]) > 0 and prompt_chunks[0][0] == tokenizer.bos_token_id ): offset = 1 input_ids.append(prompt_chunks[0][0]) for x in insert_separator(prompt_chunks, [image_token_index] * (offset + 1)): input_ids.extend(x[offset:]) if return_tensors is not None: if return_tensors == "pt": return torch.tensor(input_ids, dtype=torch.long) raise ValueError(f"Unsupported tensor type: {return_tensors}") return input_ids prompt = "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <im_start><image><im_end>\nPlease locate the tape in this image. ASSISTANT:" image = Image.open('path/to/image') image_clip = clip_image_processor.preprocess(image, return_tensors="pt")["pixel_values"][0].unsqueeze(0).cuda() image_clip = image_clip.bfloat16() input_ids = tokenizer_image_token(prompt, vsm_tokenizer, return_tensors="pt") input_ids = input_ids.unsqueeze(0).cuda() with torch.no_grad(): outputs = vsm_model.generate( images=image_clip, input_ids=input_ids, max_new_tokens=100, num_beams=1, output_hidden_states=True, return_dict_in_generate=True, )
内容的提问来源于stack exchange,提问作者Raymond Li
相关产品推荐
相关产品推荐

