You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Hugging Face tokenizer.batch_encode_plus时不同数据集列数不一致如何解决?

解决Tokenizer编码后训练/测试特征列数不一致的问题

问题现象

使用tokenizer.batch_encode_plus处理训练和测试文本时,同一Tokenizer生成的特征DataFrame列数不一致:

df_test_feats.shape
Out[2]: (2, 8)

df_train_feats.shape
Out[3]: (2, 20)

这种差异会导致传入XGBoost模型时出错,相关代码如下:

import os, sys
import pandas as pd

import torch
from transformers import AutoModel, AutoTokenizer
str_token = 'distilbert-base-uncased'


if __name__ == '__main__':
        device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
        
        model_check_point = 'distilbert-base-uncased'
        model = AutoModel.from_pretrained(model_check_point)
        tokenizer = AutoTokenizer.from_pretrained(model_check_point, add_prefix_space=True, use_fast=False)
        df_train_feats_encoded = tokenizer.batch_encode_plus(["today I went to the movies ", "today I went to the movies and had dinner at saints a new resturant in italy"], max_length=20, padding=True)
        df_train_feats = pd.DataFrame(df_train_feats_encoded['input_ids'])
        
        df_test_feats_encoded = tokenizer.batch_encode_plus(['we could not play paddal','it rain most of the afternoon'], max_length=20, padding=True)
        df_test_feats = pd.DataFrame(df_test_feats_encoded['input_ids'])

问题原因

padding=True时,Tokenizer默认采用padding='longest'策略——仅将当前批次的文本padding到该批次内最长序列的长度。训练批次里有文本被截断到max_length=20,所以训练特征列数是20;而测试批次的最长文本仅对应8个token,因此测试特征只padding到8列,最终导致两者形状不匹配。

解决方案

方案1:编码时强制padding到max_length

修改batch_encode_plus的padding参数为'max_length',这样不管当前批次文本长度如何,都会统一padding到max_length指定的长度:

# 训练集编码
df_train_feats_encoded = tokenizer.batch_encode_plus(
    ["today I went to the movies ", "today I went to the movies and had dinner at saints a new resturant in italy"],
    max_length=20,
    padding='max_length',  # 改为max_length
    truncation=True  # 显式开启截断,确保超长文本被截断到max_length
)

# 测试集编码
df_test_feats_encoded = tokenizer.batch_encode_plus(
    ['we could not play paddal','it rain most of the afternoon'],
    max_length=20,
    padding='max_length',  # 改为max_length
    truncation=True
)

修改后,训练和测试特征的input_ids都会是长度为20的列表,生成的DataFrame列数统一为20。

方案2:对已生成的测试DataFrame补全列

如果已经生成了测试集DataFrame,可以手动补充缺失的列,用DistilBERT的padding token id(默认是0)填充:

# 计算需要补充的列数
pad_cols = 20 - df_test_feats.shape[1]
# 新增列并填充0
for i in range(pad_cols):
    df_test_feats[df_test_feats.shape[1]] = 0
# 确保列索引连续(可选)
df_test_feats = df_test_feats.reindex(columns=range(20))

执行后df_test_feats.shape会变为(2,20),和训练集形状一致。

内容的提问来源于stack exchange,提问作者Sade

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 06:23:17