基于GB、树模型、随机森林的房价预测MSE过高问题求助
房价预测模型MSE过高问题排查
我用梯度提升(GB)、决策树、随机森林做Kaggle房价预测(House Prices - Advanced Regression Techniques)时,均方误差(MSE)数值始终居高不下。试过用全部变量和筛选部分变量,MSE还是很高,怀疑代码有问题;尝试过特征工程,但导致MSE恶化,暂时注释了相关代码。以下是我的代码及运行结果:
# Kaggle : House Prices - Advanced Regression Techniques import pandas as pd import numpy as np import sklearn.model_selection import matplotlib.pyplot as plt import sklearn as sk import sklearn.tree import sklearn.ensemble df = pd.read_csv('/Users/andrewhashoush/Downloads/house-prices-advanced-regression-techniques/train.csv') df_test = pd.read_csv('/Users/andrewhashoush/Downloads/house-prices-advanced-regression-techniques/test.csv') print(df.head()) print(df.info()) # selected_features = ['OverallQual', 'YearBuilt', 'TotalBsmtSF', '1stFlrSF', 'GrLivArea', # 'GarageCars', 'GarageArea', 'MSZoning', 'Neighborhood', # 'KitchenQual', 'CentralAir', 'LotArea', 'MSSubClass', 'LotFrontage', # 'Street', 'LandContour', 'Utilities', 'OverallCond', 'RoofStyle', # 'RoofMatl', 'BsmtQual','SaleCondition', 'SaleType', 'YrSold', 'MoSold', # 'PoolArea'] selected_features = [ 'LotFrontage', 'OverallQual', 'OverallCond', 'MasVnrArea', 'HalfBath', 'BedroomAbvGr', 'KitchenAbvGr', 'GarageCars', 'WoodDeckSF', 'OpenPorchSF', 'MoSold', 'YrSold', 'MSZoning', 'Alley', 'LotShape', 'LandContour', 'LotConfig', 'LandSlope', 'Neighborhood', 'Condition1', 'BldgType', 'HouseStyle', 'RoofStyle', 'Exterior1st', 'Exterior2nd', 'MasVnrType', 'ExterQual', 'ExterCond', 'Foundation', 'BsmtQual', 'BsmtCond', 'BsmtExposure', 'BsmtFinType1', 'BsmtFinType2', 'Heating', 'HeatingQC', 'CentralAir', 'Electrical', 'KitchenQual', 'Functional', 'FireplaceQu', 'GarageType', 'GarageFinish', 'GarageQual', 'GarageCond', 'PavedDrive', 'PoolQC', 'Fence', 'MiscFeature', 'SaleType', 'SaleCondition' ] # feature engineering # df['Quality_Condition'] = df['OverallQual'] * df['OverallCond'] # df['Age_at_Sale'] = df['YrSold'] - df['YearBuilt'] # selected_features += ['Quality_Condition', 'Age_at_Sale'] X = df[selected_features] print(X.head()) for column in X.columns: missing_data = df[column].isnull().sum() print(f"{column}: {missing_data}") categorical_vars = X.select_dtypes(include='object').columns.tolist() numerical_vars = X.select_dtypes(exclude='object').columns.tolist() print("Categorical Variables:", categorical_vars) print() print("Numerical Variables:", numerical_vars) #%% # filled the missing categorical values with mode for var in categorical_vars: mode_value = X[var].mode()[0] # X[var] = X[var].fillna(mode_value) # X.loc[:, var] = X.loc[:, var].fillna(mode_value) X.loc[:, var] = X[var].fillna(mode_value) # filled the missing numerical values with median for var in numerical_vars: median_value = X[var].median() X.loc[:, var] = X[var].fillna(median_value) # checking it for column in X.columns: missing_data = X[column].isnull().sum() print(f"{column}: {missing_data}") # one hot encoding X = pd.get_dummies(X, columns=categorical_vars) y = df['SalePrice'] # split the data X_train, X_val, y_train, y_val = sklearn.model_selection.train_test_split(X,y, train_size =.8, random_state= 123) dt_model = sklearn.tree.DecisionTreeRegressor(max_depth=5, random_state=123) dt_model.fit(X_train, y_train) y_pred = dt_model.predict(X_val) mse_val = np.mean((y_pred - y_val)**2) print(f"MSE: {mse_val}") # Random Forest Regressor rf_model = sklearn.ensemble.RandomForestRegressor(n_estimators=100, random_state=123) rf_model.fit(X_train, y_train) y_pred1 = rf_model.predict(X_val) mse_val1 = np.mean((y_pred1 - y_val)**2) print(f"MSE: {mse_val1}") # Gradient Boosting Regressor gb_model = sklearn.ensemble.GradientBoostingRegressor(n_estimators=100, learning_rate=0.1, random_state=123) gb_model.fit(X_train, y_train) y_pred2 = gb_model.predict(X_val) mse_val2 = np.mean((y_pred2 - y_val)**2) print(f"MSE: {mse_val2}")
运行结果:
MSE: 1451852149.6361678 MSE: 944388532.5714014 MSE: 755815420.2686024
内容的提问来源于stack exchange,提问作者James
相关产品推荐
相关产品推荐

