You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

HuggingFace Tokenizer批量处理方法咨询及报错排查

问题解答

文档描述无错误,是对输入结构的理解偏差

Tokenizer的__call__函数确实支持List[List[str]]类型输入,但该结构的含义是预分词的批量文本,而非你所理解的“批次嵌套批次”:

text(str, List[str], List[List[str]], 可选)—— 待编码的单条或批量序列。每个序列可以是字符串或字符串列表(预分词字符串)。若序列以字符串列表形式提供(预分词),必须设置is_split_into_words=True(以消除与批量序列的歧义)。

你写出的test = [test, test]是将两个批量数据再嵌套一层,这不属于Tokenizer预期的预分词结构,因此触发类型错误。

报错原因与is_split_into_words的关联

当传入List[List[str]]但设置is_split_into_words=False时,Tokenizer会将内层每个字符串视为独立待编码文本,外层列表作为批量维度,但嵌套两层的结构超出了预期输入范围——它仅接受单层批量(List[str]),或预分词的单条/批量文本(List[List[str]] + is_split_into_words=True)。

你的场景是要拆分数据避免OOM,无需嵌套列表,正确做法是将大列表拆分为多个小批量列表,分别传入Tokenizer处理:

正确分批处理示例

from transformers import AutoTokenizer
import torch

# 原始数据
test = ["hello this is a test", "that transforms a list of sentences", "into a list of list of sentences", "in order to emulate, in this case, two batches of the same lenght", "to be tokenized by the hf tokenizer for the defined model"]
tokenizer = AutoTokenizer.from_pretrained('distilbert-base-uncased-finetuned-sst-2-english')

# 手动拆分批次(示例:每2条为一个批次)
batch_size = 2
batches = [test[i:i+batch_size] for i in range(0, len(test), batch_size)]

# 分批执行分词与预测
for batch in batches:
    tokenized_batch = tokenizer(text=batch, padding="max_length", is_split_into_words=False, truncation=True, return_tensors="pt")
    # 模型预测逻辑
    # with torch.no_grad():
    #     logits = model(**tokenized_batch).logits

预分词结构的正确使用场景(仅当文本已按单词拆分时)

若你的文本已经完成预分词,可按如下方式调用:

pre_tokenized_text = [["hello", "this", "is", "a", "test"], ["that", "transforms", "a", "list", "of", "sentences"]]
tokenized = tokenizer(text=pre_tokenized_text, padding="max_length", is_split_into_words=True, truncation=True, return_tensors="pt")

这种场景才对应文档中List[List[str]]的输入要求,且必须配合is_split_into_words=True。

总结

  • 文档描述准确,List[List[str]]是预分词文本的批量输入格式,而非批次嵌套结构
  • 报错源于输入结构不符合预期,与is_split_into_words参数的正确使用直接相关
  • 解决OOM问题的正确方式是拆分原始列表为多个小批量,而非嵌套列表

内容的提问来源于stack exchange,提问作者Lucas Azevedo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 20:13:15