You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BERT Tokenizer训练分类模型预测时特征数不匹配报错如何解决

问题根本原因

你用HuggingFace Tokenizer时设置的padding=True是将序列填充到当前批次的最长长度,而非固定长度:

  • 训练批次最长序列长度为51,所以所有训练样本编码后维度都是51,模型训练时固定了输入特征数为51
  • 测试批次最长序列长度为55,所有测试样本编码后默认维度为55,你写的补全逻辑仅处理了长度不足51的样本,未对长度超过51的样本做截断,最终输入模型的特征维度为55,和模型要求的51不匹配触发报错。

修复方案

1. 训练阶段修改(推荐)

不要依赖批次动态长度,显式指定固定输入维度,修改tokenizer参数:

def tokenization_and_encoding(data,model_name,independent_col,target_col, fix_len=51):
    tokenizer = BertTokenizerFast.from_pretrained(model_name,do_lower_case=True)
    train_text=list(data[independent_col])
    train_labels=list(data[target_col])
    # 关键修改:padding设为max_length,直接输出固定长度的编码,不需要后续手动处理
    train_encodings = tokenizer(train_text, truncation=True, padding='max_length', max_length=fix_len)
    train_encodings=train_encodings['input_ids']
    return train_encodings,train_labels

# 调用时指定固定长度
train_encodings,train_labels = tokenization_and_encoding(data, model_name, 'Description', 'Target', fix_len=51)

额外隐藏bug提醒:你代码中筛选非空值时用的列名是Outage Description,但前面选列时用的是Description,列名不一致会导致空值过滤失效,建议统一列名。

2. 预测阶段修改

和训练阶段保持一致的tokenizer参数,不需要手动写循环补全/截断逻辑,直接输出固定长度编码:

# 编码测试文本,和训练参数完全对齐
test_encodings = tokenizer(test_text, truncation=True, padding='max_length', max_length=51)
test_encodings=test_encodings['input_ids']

# 加载模型后直接预测即可
predict_svc_test_pred_bbc = pd.DataFrame(Pickled_svc_Model.predict(test_encodings))

如果不想重新训练模型,也可以在当前预测代码里补充截断逻辑,修改你现有的补全循环:

for enc in test_encodings:    
    # 先截断超过51的部分
    if len(enc) > 51:
        enc = enc[:51]
    # 再补全不足51的部分
    else:
        enc.extend([0]*(51 - len(enc)))

内容的提问来源于stack exchange,提问作者Aditya sharma

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 15:54:04