You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Sklearn的Airbnb房价线性回归拟合不佳,求代码优化帮助

Airbnb房价预测:线性回归模型优化方案

你的模型出现大量数据偏离预测线的核心原因是特征维度不足,且未针对房价预测的特性做针对性处理。以下是具体优化步骤:

1. 补充关键预测特征

房价受多种因素影响,仅用room_type和neighbourhood远远不够。建议加入Airbnb数据集中常见的高相关性特征:

# 扩展特征列表,根据你的数据集调整可用特征
X = df2[['room_type', 'neighbourhood', 'accommodates', 'bedrooms', 
         'bathrooms', 'review_scores_rating', 'minimum_nights', 'availability_365']]

2. 标准化数值特征

线性回归对特征尺度敏感,数值特征(如accommodates、bedrooms)需要标准化处理,避免尺度差异主导模型训练:

from sklearn.preprocessing import StandardScaler

# 分离数值与分类特征
numeric_cols = ['accommodates', 'bedrooms', 'bathrooms', 'review_scores_rating']
categorical_cols = ['room_type', 'neighbourhood']

# 标准化数值特征
scaler = StandardScaler()
scaled_numeric = scaler.fit_transform(df2[numeric_cols])
scaled_numeric_df = pd.DataFrame(scaled_numeric, columns=numeric_cols, index=df2.index)

# 编码分类特征(保持原有逻辑)
encoded_categorical = pd.get_dummies(df2[categorical_cols], drop_first=True)

# 合并处理后的特征
X = pd.concat([scaled_numeric_df, encoded_categorical], axis=1)

3. 改用正则化线性模型

普通线性回归容易受异常值或噪声影响,加入正则项约束系数,提升模型泛化能力:

from sklearn.linear_model import Ridge

# Ridge回归(L2正则),alpha为正则强度,可通过网格搜索调优
model = Ridge(alpha=1.0)
model.fit(X_train, y_train)

# 若想筛选重要特征,可改用Lasso回归(L1正则)
# from sklearn.linear_model import Lasso
# model = Lasso(alpha=0.5)
# model.fit(X_train, y_train)

4. 修正目标变量分布

房价通常呈右偏分布,违背线性回归的正态性假设,对目标变量做对数转换:

import numpy as np

# 用log1p避免0值问题,预测后需用expm1转回原始尺度
y = np.log1p(df2['price'])

# 训练模型后,预测值转换回原始房价
predictions = np.expm1(model.predict(X_test))

5. 添加交互特征

不同街区的房间类型可能存在价格协同效应,构造交互特征捕捉这类关系:

# 生成房间类型与街区的交互特征
df2['room_type_neighbourhood'] = df2['room_type'] + '_' + df2['neighbourhood']

# 将交互特征加入分类特征列表重新编码
categorical_cols = ['room_type', 'neighbourhood', 'room_type_neighbourhood']
encoded_categorical = pd.get_dummies(df2[categorical_cols], drop_first=True)

6. 量化模型评估

仅靠回归图不够,用量化指标评估模型效果,方便对比优化前后的提升:

from sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score

mae = mean_absolute_error(y_test, predictions)
mse = mean_squared_error(y_test, predictions)
r2 = r2_score(y_test, predictions)

print(f"平均绝对误差(MAE): {mae:.2f}")
print(f"均方误差(MSE): {mse:.2f}")
print(f"决定系数(R²): {r2:.2f}")

内容的提问来源于stack exchange,提问作者Nico

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 13:02:47