You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

ViT模型:Hugging Face Pipeline与手动提取嵌入结果不一致问题

ViT嵌入提取:Pipeline与手动流程差异问题解答

问题背景

我使用transformers库的ViTForImageClassification模型(google/vit-base-patch16-224)提取图像嵌入时,发现用预构建pipeline函数和手动提取流程得到的嵌入结果不一致,但二者目标模型架构和层级一致。现提出以下问题:

  1. pipeline与手动方式在图像预处理或特征提取上是否存在固有差异,导致该结果不一致?
  2. 如何确保两种方式提取的嵌入具备可比性?
  3. pipeline函数是否存在未在手动方式中显式设置的额外配置,影响输出结果?

测试代码

# %% data
import requests
import torchvision
from PIL import Image

# url = "https://huggingface.co/datasets/Narsil/image_dummy/raw/main/parrots.png"

url = 'http://images.cocodataset.org/val2017/000000039769.jpg'
pil_image = Image.open(requests.get(url, stream=True).raw)

embeddings = {}

# %% APPROACH 1
# note: prints warning: Some weights of ViTModel were not initialized from the model checkpoint at google/vit-base-patch16-224 and are newly initialized ['vit.pooler.dense.bias', 'vit.pooler.dense.weight']
from transformers import pipeline

extractor = pipeline(model="google/vit-base-patch16-224", task="image-feature-extraction")
y = extractor(pil_image, return_tensors=True)

embeddings['transformers_pipeline'] = y

# %% APPROACH 2
from transformers import ViTImageProcessor, ViTForImageClassification

processor = ViTImageProcessor.from_pretrained('google/vit-base-patch16-224')
model = ViTForImageClassification.from_pretrained('google/vit-base-patch16-224',
                                                  output_attentions=False,
                                                  output_hidden_states=True)  # Enable outputting hidden

inputs = processor(images=pil_image, return_tensors="pt")
outputs = model(**inputs)
y = outputs.hidden_states[-1]

embeddings['transformers_base'] = y

问题1:差异根源

核心差异来自模型类型和特征提取位置的不一致,预处理环节基本无差异:

  • 模型类型不同:pipeline执行image-feature-extraction任务时,会自动加载ViTModel(纯特征提取模型,带pooler层);而手动流程用的是ViTForImageClassification(带分类头的任务模型)。
  • 特征提取位置不同:
    • pipeline默认输出的是ViTModel的pooler_output(对最后一层的<[BOS_never_used_51bce0c785ca2f68081bfa7d91973934]> token做线性变换后的结果);
    • 手动流程取的是ViTForImageClassification的hidden_states[-1],即最后一层所有token的原始嵌入(包含<[BOS_never_used_51bce0c785ca2f68081bfa7d91973934]>但未经过pooler处理)。

问题2:确保嵌入可比性的方法

要让两种方式输出一致,需要对齐模型类型和特征提取逻辑:

方法一:手动流程改用ViTModel

和pipeline加载相同的模型类型,提取对应特征:

from transformers import ViTImageProcessor, ViTModel

processor = ViTImageProcessor.from_pretrained('google/vit-base-patch16-224')
# 加载ViTModel,和pipeline使用的模型一致
model = ViTModel.from_pretrained('google/vit-base-patch16-224', output_hidden_states=True)
# 显式切换到eval模式,避免训练层干扰
model.eval()

inputs = processor(images=pil_image, return_tensors="pt")
outputs = model(**inputs)
# 取pooler输出,和pipeline默认行为对齐
y = outputs.pooler_output
# 或者直接取最后一层的<[BOS_never_used_51bce0c785ca2f68081bfa7d91973934]> token嵌入(与pooler输出逻辑一致)
# y = outputs.last_hidden_state[:, 0, :]

方法二:修改pipeline的输出配置

如果坚持使用ViTForImageClassification,可以让pipeline输出所有token的原始嵌入,和手动流程对齐:

# 设置return_all_tokens=True,输出最后一层所有token的嵌入
extractor = pipeline(model="google/vit-base-patch16-224", task="image-feature-extraction", return_all_tokens=True)
y = extractor(pil_image, return_tensors=True)
# 此时y与手动流程的outputs.hidden_states[-1]完全一致

问题3:pipeline的额外默认配置

pipeline存在几个手动流程可能未显式设置的默认行为:

  • 自动切换模型到eval模式:pipeline初始化时会自动调用model.eval(),关闭dropout、batch norm等训练相关层;手动流程如果没显式执行这一步,模型处于train模式,会导致输出随机性差异。
  • 默认特征输出策略:image-feature-extraction pipeline默认输出pooler_output,而非所有token的嵌入;手动流程如果没指定对应输出位置,就会出现差异。
  • 自动设备分配:pipeline会自动将模型和输入移到可用GPU(若存在);手动流程若未显式指定设备,会在CPU运行,但精度差异可忽略,不影响结果一致性。
  • pooler层权重初始化:pipeline加载ViTModel时,会初始化预训练checkpoint中没有的pooler层权重(也就是你看到的警告);而手动流程用ViTForImageClassification时不会加载pooler层,因此无此警告。

内容的提问来源于stack exchange,提问作者martinelliadr

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 16:33:16