使用Hugging Face tokenizer.batch_encode_plus时不同数据集列数不一致如何解决?
解决Tokenizer编码后训练/测试特征列数不一致的问题
问题现象
使用tokenizer.batch_encode_plus处理训练和测试文本时,同一Tokenizer生成的特征DataFrame列数不一致:
df_test_feats.shape Out[2]: (2, 8) df_train_feats.shape Out[3]: (2, 20)
这种差异会导致传入XGBoost模型时出错,相关代码如下:
import os, sys import pandas as pd import torch from transformers import AutoModel, AutoTokenizer str_token = 'distilbert-base-uncased' if __name__ == '__main__': device = torch.device('cuda' if torch.cuda.is_available() else 'cpu') model_check_point = 'distilbert-base-uncased' model = AutoModel.from_pretrained(model_check_point) tokenizer = AutoTokenizer.from_pretrained(model_check_point, add_prefix_space=True, use_fast=False) df_train_feats_encoded = tokenizer.batch_encode_plus(["today I went to the movies ", "today I went to the movies and had dinner at saints a new resturant in italy"], max_length=20, padding=True) df_train_feats = pd.DataFrame(df_train_feats_encoded['input_ids']) df_test_feats_encoded = tokenizer.batch_encode_plus(['we could not play paddal','it rain most of the afternoon'], max_length=20, padding=True) df_test_feats = pd.DataFrame(df_test_feats_encoded['input_ids'])
问题原因
padding=True时,Tokenizer默认采用padding='longest'策略——仅将当前批次的文本padding到该批次内最长序列的长度。训练批次里有文本被截断到max_length=20,所以训练特征列数是20;而测试批次的最长文本仅对应8个token,因此测试特征只padding到8列,最终导致两者形状不匹配。
解决方案
方案1:编码时强制padding到max_length
修改batch_encode_plus的padding参数为'max_length',这样不管当前批次文本长度如何,都会统一padding到max_length指定的长度:
# 训练集编码 df_train_feats_encoded = tokenizer.batch_encode_plus( ["today I went to the movies ", "today I went to the movies and had dinner at saints a new resturant in italy"], max_length=20, padding='max_length', # 改为max_length truncation=True # 显式开启截断,确保超长文本被截断到max_length ) # 测试集编码 df_test_feats_encoded = tokenizer.batch_encode_plus( ['we could not play paddal','it rain most of the afternoon'], max_length=20, padding='max_length', # 改为max_length truncation=True )
修改后,训练和测试特征的input_ids都会是长度为20的列表,生成的DataFrame列数统一为20。
方案2:对已生成的测试DataFrame补全列
如果已经生成了测试集DataFrame,可以手动补充缺失的列,用DistilBERT的padding token id(默认是0)填充:
# 计算需要补充的列数 pad_cols = 20 - df_test_feats.shape[1] # 新增列并填充0 for i in range(pad_cols): df_test_feats[df_test_feats.shape[1]] = 0 # 确保列索引连续(可选) df_test_feats = df_test_feats.reindex(columns=range(20))
执行后df_test_feats.shape会变为(2,20),和训练集形状一致。
内容的提问来源于stack exchange,提问作者Sade
相关产品推荐
相关产品推荐

