You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

模型训练预测报错ValueError:特征名称与拟合时不匹配的解决方法

解决模型训练预测时的特征名称不匹配问题

问题背景

训练模型时已通过删除InvoiceDate、BillingAddress等含字符串的列解决了字符串转浮点的报错,但执行预测时出现以下错误:

ValueError: The feature names should match those that were passed during fit. Feature names unseen at fit time:

相关代码片段:

X_train = X_train.drop(columns=['InvoiceDate', "BillingAddress", "BillingCity", "BillingState", "BillingCountry", "BillingPostalCode", "Rowversion_x", "Rowversion_y", "Rowversion", "Name", "Composer"], axis=1)

print(f'Tipo X_train: {type(X_train)} Tipo y_train: {type(y_train)}')
clf.fit(X_train, y_train)
y_pred = clf.predict(X_test)

accuracy = accuracy_score(y_test, y_pred)
print("Accuracy:", accuracy)

mining_data.to_csv('mining_table.csv', index=False)

解决办法

  • 同步处理X_test的特征列
    你仅对X_train执行了列删除操作,但X_test保留了原始全部特征列,导致模型预测时的特征集与训练时不匹配。必须对X_test执行完全相同的列删除操作:

    # 用统一的列名列表处理,避免手动输入错误
    drop_cols = ['InvoiceDate', "BillingAddress", "BillingCity", "BillingState", "BillingCountry", "BillingPostalCode", "Rowversion_x", "Rowversion_y", "Rowversion", "Name", "Composer"]
    X_train = X_train.drop(columns=drop_cols, axis=1)
    X_test = X_test.drop(columns=drop_cols, axis=1)
    
  • 验证特征列一致性
    在训练和预测前,检查X_train和X_test的特征列是否完全一致:

    # 打印特征列对比
    print("X_train特征列:", X_train.columns.tolist())
    print("X_test特征列:", X_test.columns.tolist())
    # 直接判断一致性
    print("特征列是否一致:", set(X_train.columns) == set(X_test.columns))
    

    若输出不一致,排查是否有其他代码修改了某一方的特征列。

  • 用Pipeline封装预处理流程(可选)
    为避免后续重复手动处理特征列的问题,可使用sklearn.pipeline.Pipeline将预处理与模型训练封装在一起,确保训练和预测的预处理逻辑完全一致:

    from sklearn.pipeline import Pipeline
    from sklearn.base import BaseEstimator, TransformerMixin
    
    # 自定义删除列的变换器
    class DropColumns(BaseEstimator, TransformerMixin):
        def __init__(self, cols):
            self.cols = cols
        def fit(self, X, y=None):
            return self
        def transform(self, X):
            return X.drop(columns=self.cols, axis=1)
    
    # 构建Pipeline
    drop_cols = ['InvoiceDate', "BillingAddress", "BillingCity", "BillingState", "BillingCountry", "BillingPostalCode", "Rowversion_x", "Rowversion_y", "Rowversion", "Name", "Composer"]
    pipeline = Pipeline([
        ('drop_cols', DropColumns(cols=drop_cols)),
        ('classifier', clf)
    ])
    
    # 训练和预测
    pipeline.fit(X_train, y_train)
    y_pred = pipeline.predict(X_test)
    

内容的提问来源于stack exchange,提问作者information

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 02:43:25