You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

同一LLaVA模型在不同Transformers版本中的分词与生成差异求助

问题描述

我下载了一个基于LLaVA的自定义旧模型,该模型原本适配Transformers 4.31.0版本。由于需要搭配使用适配Transformers 4.53.1版本的Qwen模型,我升级了Transformers版本,此后该LLaVA模型无法正常运行。经排查发现,即使模型及配置完全相同,同一prompt在不同Transformers版本下的分词与生成行为仍存在差异。

差异示例

测试Prompt

"A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <im_start><image><im_end>\nPlease locate the tape in this image. ASSISTANT:"

分词结果差异

  • Transformers 4.31.0分词结果:
    Tokenized 4.31.0
  • Transformers 4.53.1分词结果:
    Tokenized 4.53.1

模型输出差异

  • Transformers 4.31.0输出:
    outputs 4.31.0
  • Transformers 4.53.1输出:
    outputs 4.53.1

疑问

我想了解为何会出现这种模型行为差异,以及是否可以让该模型恢复在Transformers 4.31.0版本下的运行行为。

示例代码

from transformers import AutoTokenizer, CLIPImageProcessor
import torch
from PIL import Image

kwargs = {}
kwargs['torch_dtype'] = torch.bfloat16
kwargs['device_map'] = 'cuda'
kwargs['is_eval'] = True
vsm_tokenizer = AutoTokenizer.from_pretrained(
    'craigwu/seal_vsm_7b',
    cache_dir=None,
    model_max_length=512,
    padding_side="right",
    use_fast=False
)

vsm_tokenizer.pad_token = vsm_tokenizer.unk_token
loc_token_idx = vsm_tokenizer("[LOC]", add_special_tokens=False).input_ids[0]
vsm_model = VSMForCausalLM.from_pretrained(
    'craigwu/seal_vsm_7b', low_cpu_mem_usage=True, vision_tower='openai/clip-vit-large-patch14',
    loc_token_idx=loc_token_idx, **kwargs
)
vsm_model.get_model().initialize_vision_modules(vsm_model.get_model().config)
vsm_model.eval()
clip_image_processor = CLIPImageProcessor.from_pretrained('openai/clip-vit-large-patch14')


def tokenizer_image_token(
    prompt, tokenizer, image_token_index=IMAGE_TOKEN_INDEX, return_tensors=None
):
    prompt_chunks = [tokenizer(chunk).input_ids for chunk in prompt.split("<image>")]

    def insert_separator(X, sep):
        return [ele for sublist in zip(X, [sep] * len(X)) for ele in sublist][:-1]

    input_ids = []
    offset = 0
    if (
        len(prompt_chunks) > 0
        and len(prompt_chunks[0]) > 0
        and prompt_chunks[0][0] == tokenizer.bos_token_id
    ):
        offset = 1
        input_ids.append(prompt_chunks[0][0])

    for x in insert_separator(prompt_chunks, [image_token_index] * (offset + 1)):
        input_ids.extend(x[offset:])

    if return_tensors is not None:
        if return_tensors == "pt":
            return torch.tensor(input_ids, dtype=torch.long)
        raise ValueError(f"Unsupported tensor type: {return_tensors}")
    return input_ids

prompt = "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <im_start><image><im_end>\nPlease locate the tape in this image. ASSISTANT:"
image = Image.open('path/to/image')

image_clip = clip_image_processor.preprocess(image, return_tensors="pt")["pixel_values"][0].unsqueeze(0).cuda()

image_clip = image_clip.bfloat16()
input_ids = tokenizer_image_token(prompt, vsm_tokenizer, return_tensors="pt")
input_ids = input_ids.unsqueeze(0).cuda()

with torch.no_grad():
    outputs = vsm_model.generate(
        images=image_clip,
        input_ids=input_ids,
        max_new_tokens=100,
        num_beams=1,
        output_hidden_states=True,
        return_dict_in_generate=True,
    )

内容的提问来源于stack exchange,提问作者Raymond Li

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 17:37:04