如何基于训练好的回归模型从test.csv预测房屋售价?
如何用训练好的模型预测test.csv中的房屋售价
你的代码已经完成了模型训练,但要注意数据预处理的一致性——测试集必须和训练集用完全相同的预处理逻辑(比如LabelEncoder、OneHotEncoder、标准化的参数都要复用训练集的,不能重新fit),否则会导致预测结果出错。以下是完整的修改步骤和代码:
核心修改点
- 保存训练时用到的所有预处理组件(LabelEncoder、OneHotEncoder、标准化器),用于测试集的统一处理
- 用训练集的规则预处理test.csv,确保特征维度和分布与训练集一致
- 预测后将标准化的结果反转换为真实售价,并保存为可提交的格式
修改后的完整代码
import numpy as np import pandas as pd from sklearn.preprocessing import LabelEncoder, OneHotEncoder, StandardScaler from sklearn.ensemble import RandomForestRegressor from sklearn.metrics import r2_score from sklearn.model_selection import train_test_split # ---------------------- 训练阶段:保存预处理组件与模型 ---------------------- # 加载训练数据 data = pd.read_csv("train.csv") categorical_cols = ['MSZoning','Street','Alley','LotShape','LandContour','Utilities','LotConfig','LandSlope','Neighborhood','Condition1','Condition2','BldgType','HouseStyle','RoofStyle','RoofMatl','Exterior1st','Exterior2nd','MasVnrType','ExterQual','ExterCond','Foundation','BsmtQual','BsmtCond','BsmtExposure','BsmtFinType1','BsmtFinType2','Heating','HeatingQC','CentralAir','Electrical','KitchenQual','Functional','FireplaceQu','GarageType','GarageFinish','GarageQual','GarageCond','PavedDrive','PoolQC','Fence','MiscFeature','SaleType','SaleCondition'] # 为每个分类列单独保存LabelEncoder,避免编码冲突 label_encoders = {} for col in categorical_cols: le = LabelEncoder() data[col] = data[col].fillna('Unknown') # 统一处理缺失分类值 data[col] = le.fit_transform(data[col]) label_encoders[col] = le # 保存OneHotEncoder ohe = OneHotEncoder(sparse_output=False, drop='first') ohe.fit(data[categorical_cols]) array_hot_encoded = ohe.transform(data[categorical_cols]) data_hot_encoded = pd.DataFrame(array_hot_encoded, columns=ohe.get_feature_names_out(categorical_cols), index=data.index) # 合并特征并处理缺失值 data_other_cols = data.drop(columns=categorical_cols + ['Id']) data_out = pd.concat([data_hot_encoded, data_other_cols], axis=1) data_out = data_out.fillna(method="bfill").dropna() # 分离特征与目标变量,保存标准化器 X = data_out.drop(columns=['SalePrice']) y = data_out['SalePrice'] scaler_X = StandardScaler() X_scaled = scaler_X.fit_transform(X) scaler_y = StandardScaler() y_scaled = scaler_y.fit_transform(y.values.reshape(-1, 1)).squeeze() # 划分验证集并训练模型 X_train, X_test, y_train, y_test = train_test_split(X_scaled, y_scaled, test_size=0.2) model = RandomForestRegressor() model.fit(X_train, y_train) # 输出验证集R2分数 y_pred = model.predict(X_test) print(f"验证集R2 Score: {r2_score(y_test, y_pred)}") # ---------------------- 预测阶段:处理test.csv并生成结果 ---------------------- # 加载测试数据,保存Id用于结果输出 test_data = pd.read_csv("test.csv") test_ids = test_data['Id'] # 复用训练集的LabelEncoder处理测试集分类特征 for col in categorical_cols: test_data[col] = test_data[col].fillna('Unknown') test_data[col] = label_encoders[col].transform(test_data[col]) # 复用训练集的OneHotEncoder编码 test_hot_encoded = ohe.transform(test_data[categorical_cols]) test_hot_encoded_df = pd.DataFrame(test_hot_encoded, columns=ohe.get_feature_names_out(categorical_cols), index=test_data.index) # 合并特征并处理缺失值 test_other_cols = test_data.drop(columns=categorical_cols + ['Id']) test_out = pd.concat([test_hot_encoded_df, test_other_cols], axis=1) test_out = test_out.fillna(method="bfill").dropna() # 复用训练集的标准化器处理特征 test_X_scaled = scaler_X.transform(test_out) # 预测并反标准化得到真实售价 test_y_scaled = model.predict(test_X_scaled) test_y_pred = scaler_y.inverse_transform(test_y_scaled.reshape(-1, 1)).squeeze() # 保存预测结果为CSV result_df = pd.DataFrame({'Id': test_ids[:len(test_y_pred)], 'SalePrice': test_y_pred}) result_df.to_csv("submission.csv", index=False) print("预测结果已保存为submission.csv")
关键注意事项
- 预处理一致性:所有编码器和标准化器必须复用训练集的参数,绝对不能在测试集上重新
fit,否则会破坏特征分布一致性 - 缺失值处理:测试集的缺失值处理逻辑要和训练集完全对齐,比如统一用
bfill填充、未知分类用Unknown替代 - 反标准化:如果训练时对目标变量做了标准化,预测后必须用
scaler_y.inverse_transform转换回真实的售价范围
内容的提问来源于stack exchange,提问作者Athira Jayaram
相关产品推荐
相关产品推荐

