HuggingFace Tokenizer批量处理方法咨询及报错排查
问题解答
文档描述无错误,是对输入结构的理解偏差
Tokenizer的__call__函数确实支持List[List[str]]类型输入,但该结构的含义是预分词的批量文本,而非你所理解的“批次嵌套批次”:
text(str, List[str], List[List[str]], 可选)—— 待编码的单条或批量序列。每个序列可以是字符串或字符串列表(预分词字符串)。若序列以字符串列表形式提供(预分词),必须设置is_split_into_words=True(以消除与批量序列的歧义)。
你写出的test = [test, test]是将两个批量数据再嵌套一层,这不属于Tokenizer预期的预分词结构,因此触发类型错误。
报错原因与is_split_into_words的关联
当传入List[List[str]]但设置is_split_into_words=False时,Tokenizer会将内层每个字符串视为独立待编码文本,外层列表作为批量维度,但嵌套两层的结构超出了预期输入范围——它仅接受单层批量(List[str]),或预分词的单条/批量文本(List[List[str]] + is_split_into_words=True)。
你的场景是要拆分数据避免OOM,无需嵌套列表,正确做法是将大列表拆分为多个小批量列表,分别传入Tokenizer处理:
正确分批处理示例
from transformers import AutoTokenizer import torch # 原始数据 test = ["hello this is a test", "that transforms a list of sentences", "into a list of list of sentences", "in order to emulate, in this case, two batches of the same lenght", "to be tokenized by the hf tokenizer for the defined model"] tokenizer = AutoTokenizer.from_pretrained('distilbert-base-uncased-finetuned-sst-2-english') # 手动拆分批次(示例:每2条为一个批次) batch_size = 2 batches = [test[i:i+batch_size] for i in range(0, len(test), batch_size)] # 分批执行分词与预测 for batch in batches: tokenized_batch = tokenizer(text=batch, padding="max_length", is_split_into_words=False, truncation=True, return_tensors="pt") # 模型预测逻辑 # with torch.no_grad(): # logits = model(**tokenized_batch).logits
预分词结构的正确使用场景(仅当文本已按单词拆分时)
若你的文本已经完成预分词,可按如下方式调用:
pre_tokenized_text = [["hello", "this", "is", "a", "test"], ["that", "transforms", "a", "list", "of", "sentences"]] tokenized = tokenizer(text=pre_tokenized_text, padding="max_length", is_split_into_words=True, truncation=True, return_tensors="pt")
这种场景才对应文档中List[List[str]]的输入要求,且必须配合is_split_into_words=True。
总结
- 文档描述准确,
List[List[str]]是预分词文本的批量输入格式,而非批次嵌套结构 - 报错源于输入结构不符合预期,与
is_split_into_words参数的正确使用直接相关 - 解决OOM问题的正确方式是拆分原始列表为多个小批量,而非嵌套列表
内容的提问来源于stack exchange,提问作者Lucas Azevedo
相关产品推荐
相关产品推荐

