You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于训练好的回归模型从test.csv预测房屋售价?

如何用训练好的模型预测test.csv中的房屋售价

你的代码已经完成了模型训练,但要注意数据预处理的一致性——测试集必须和训练集用完全相同的预处理逻辑(比如LabelEncoder、OneHotEncoder、标准化的参数都要复用训练集的,不能重新fit),否则会导致预测结果出错。以下是完整的修改步骤和代码:

核心修改点

  1. 保存训练时用到的所有预处理组件(LabelEncoder、OneHotEncoder、标准化器),用于测试集的统一处理
  2. 用训练集的规则预处理test.csv,确保特征维度和分布与训练集一致
  3. 预测后将标准化的结果反转换为真实售价,并保存为可提交的格式

修改后的完整代码

import numpy as np
import pandas as pd
from sklearn.preprocessing import LabelEncoder, OneHotEncoder, StandardScaler
from sklearn.ensemble import RandomForestRegressor
from sklearn.metrics import r2_score
from sklearn.model_selection import train_test_split

# ---------------------- 训练阶段:保存预处理组件与模型 ----------------------
# 加载训练数据
data = pd.read_csv("train.csv")
categorical_cols = ['MSZoning','Street','Alley','LotShape','LandContour','Utilities','LotConfig','LandSlope','Neighborhood','Condition1','Condition2','BldgType','HouseStyle','RoofStyle','RoofMatl','Exterior1st','Exterior2nd','MasVnrType','ExterQual','ExterCond','Foundation','BsmtQual','BsmtCond','BsmtExposure','BsmtFinType1','BsmtFinType2','Heating','HeatingQC','CentralAir','Electrical','KitchenQual','Functional','FireplaceQu','GarageType','GarageFinish','GarageQual','GarageCond','PavedDrive','PoolQC','Fence','MiscFeature','SaleType','SaleCondition']

# 为每个分类列单独保存LabelEncoder,避免编码冲突
label_encoders = {}
for col in categorical_cols:
    le = LabelEncoder()
    data[col] = data[col].fillna('Unknown')  # 统一处理缺失分类值
    data[col] = le.fit_transform(data[col])
    label_encoders[col] = le

# 保存OneHotEncoder
ohe = OneHotEncoder(sparse_output=False, drop='first')
ohe.fit(data[categorical_cols])
array_hot_encoded = ohe.transform(data[categorical_cols])
data_hot_encoded = pd.DataFrame(array_hot_encoded, columns=ohe.get_feature_names_out(categorical_cols), index=data.index)

# 合并特征并处理缺失值
data_other_cols = data.drop(columns=categorical_cols + ['Id'])
data_out = pd.concat([data_hot_encoded, data_other_cols], axis=1)
data_out = data_out.fillna(method="bfill").dropna()

# 分离特征与目标变量,保存标准化器
X = data_out.drop(columns=['SalePrice'])
y = data_out['SalePrice']

scaler_X = StandardScaler()
X_scaled = scaler_X.fit_transform(X)

scaler_y = StandardScaler()
y_scaled = scaler_y.fit_transform(y.values.reshape(-1, 1)).squeeze()

# 划分验证集并训练模型
X_train, X_test, y_train, y_test = train_test_split(X_scaled, y_scaled, test_size=0.2)
model = RandomForestRegressor()
model.fit(X_train, y_train)

# 输出验证集R2分数
y_pred = model.predict(X_test)
print(f"验证集R2 Score: {r2_score(y_test, y_pred)}")

# ---------------------- 预测阶段:处理test.csv并生成结果 ----------------------
# 加载测试数据,保存Id用于结果输出
test_data = pd.read_csv("test.csv")
test_ids = test_data['Id']

# 复用训练集的LabelEncoder处理测试集分类特征
for col in categorical_cols:
    test_data[col] = test_data[col].fillna('Unknown')
    test_data[col] = label_encoders[col].transform(test_data[col])

# 复用训练集的OneHotEncoder编码
test_hot_encoded = ohe.transform(test_data[categorical_cols])
test_hot_encoded_df = pd.DataFrame(test_hot_encoded, columns=ohe.get_feature_names_out(categorical_cols), index=test_data.index)

# 合并特征并处理缺失值
test_other_cols = test_data.drop(columns=categorical_cols + ['Id'])
test_out = pd.concat([test_hot_encoded_df, test_other_cols], axis=1)
test_out = test_out.fillna(method="bfill").dropna()

# 复用训练集的标准化器处理特征
test_X_scaled = scaler_X.transform(test_out)

# 预测并反标准化得到真实售价
test_y_scaled = model.predict(test_X_scaled)
test_y_pred = scaler_y.inverse_transform(test_y_scaled.reshape(-1, 1)).squeeze()

# 保存预测结果为CSV
result_df = pd.DataFrame({'Id': test_ids[:len(test_y_pred)], 'SalePrice': test_y_pred})
result_df.to_csv("submission.csv", index=False)
print("预测结果已保存为submission.csv")

关键注意事项

  • 预处理一致性:所有编码器和标准化器必须复用训练集的参数,绝对不能在测试集上重新fit,否则会破坏特征分布一致性
  • 缺失值处理:测试集的缺失值处理逻辑要和训练集完全对齐,比如统一用bfill填充、未知分类用Unknown替代
  • 反标准化:如果训练时对目标变量做了标准化,预测后必须用scaler_y.inverse_transform转换回真实的售价范围

内容的提问来源于stack exchange,提问作者Athira Jayaram

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 07:36:20