You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

XGBRegressor拟合效果极差求助:自变量为分类变量,因变量连续

XGBRegressor拟合回归模型效果极差的排查与优化建议

问题背景

我是建模新手,首次用XGBRegressor拟合数据,效果极差,拟合图已展示。

数据处理流程

  • 下载公开数据集后完成大量清洗:移除零值、填充空值、删除冗余项、清理无意义极值与负值、合并类别过多的列,清洗后编码前的数据已展示。
  • 特征与目标:除目标变量value_per_ton(数值型,回归任务)外,大部分特征为分类值。
  • 编码处理:对Year和Month做Label Encoding,其余分类特征做One-Hot Encoding;用方差膨胀因子(VIF)检查共线性,所有VIF<5(多数1-3),共线性无问题。
  • 代码实现:
import pandas as pd
import numpy as np 
import matplotlib.pylab as plt
%matplotlib inline

from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_absolute_error
from sklearn import preprocessing
from sklearn.preprocessing import LabelEncoder

from xgboost import XGBRegressor
import requests

import os
for dirname, _, filenames in os.walk('/kaggle/input'):
    for filename in filenames:
        print(os.path.join(dirname, filename))

url="/kaggle/input/uk-fleets-simplified/Fleet_Output.csv"
fleetUK=pd.read_csv(url)

# Label encode the year and month columns
le = preprocessing.LabelEncoder()
fleetUK['year'] = le.fit_transform(fleetUK.year.values)
fleetUK['month'] = le.fit_transform(fleetUK.month.values)

# One-Hot Encode the categorical feature columns
fleetUK = pd.get_dummies(data=fleetUK, columns = ['port_nationality', 'length_group', 'gear_category', 'species'], drop_first=True)

# Separate X and Y
X=fleetUK.copy()
X.dropna(axis=0, subset=['value_per_ton'], inplace=True)
y = X.value_per_ton             
X.drop(['value_per_ton'], axis=1, inplace=True)

# Transforming Y to be normally distributed as it's highly skewed initially. This may not be necessary for XGBoost.
y = np.log(y)

# Train Test Splitting
X_train, X_test, y_train, y_test = train_test_split(X, y, train_size=0.7, test_size=0.3, random_state = 22)

# Defining XGBoost Regression Model and fitting on training dataset
my_model_1 = XGBRegressor(random_state = 22)
my_model_1.fit(X_train, y_train)

# Making prediction on test set
pred_1 = my_model_1.predict(X_test)

# Calculate error
mae_1 = mean_absolute_error(pred_1, y_test)
print("Mean Absolute Error:" , mae_1)

# Plotting results 
plt.figure(figsize=(10,10))
plt.scatter(y_test, pred_1, c='crimson')
plt.yscale('log')
plt.xscale('log')

p1 = max(max(pred_1), max(y_test))
p2 = min(min(pred_1), min(y_test))
plt.plot([p1, p2], [p1, p2], 'b-')
plt.xlabel('True Values', fontsize=15)
plt.ylabel('Predictions', fontsize=15)
plt.axis('equal')
plt.show()

已尝试的调整

  • 因原始目标变量分布极度偏斜,用np.log(y)转换目标值
  • 调整XGBRegressor的输入参数
  • 调整训练测试拆分比例
  • 尚未尝试对Month做循环编码,但认为这不是核心问题

以上调整均未带来明显改善,怀疑数据本身存在问题,想知道哪些环节可以调整提升拟合效果,或者XGBoost是否不适合该场景?


排查与优化建议

1. 先验证基准模型性能

先不用XGB,用简单线性回归、决策树回归这类基准模型跑一遍,对比MAE:

from sklearn.linear_model import LinearRegression
from sklearn.tree import DecisionTreeRegressor

# 线性回归基准
lr_model = LinearRegression()
lr_model.fit(X_train, y_train)
lr_pred = lr_model.predict(X_test)
print("Linear Regression MAE:", mean_absolute_error(lr_pred, y_test))

# 决策树基准
dt_model = DecisionTreeRegressor(random_state=22)
dt_model.fit(X_train, y_train)
dt_pred = dt_model.predict(X_test)
print("Decision Tree MAE:", mean_absolute_error(dt_pred, y_test))

如果基准模型效果也差,说明问题大概率在数据/特征层面,而非XGB参数;如果基准模型效果正常,再聚焦XGB的调参。

2. 特征与数据的深层排查

  • 目标变量转换的合理性:虽然XGB对分布不敏感,但np.log(y)可能导致部分极端值被过度压缩,可尝试np.sqrt(y)或分位数转换(如QuantileTransformer),或者直接用原始Y训练对比效果。
  • 分类特征的编码问题:
    • Year用Label Encoding不合理,直接保留原始年份数值即可;Month应做循环编码(如sin/cos转换),因为12月和1月是连续的,Label Encoding会拉远两者的距离:
      # Month循环编码
      fleetUK['month_sin'] = np.sin(2 * np.pi * fleetUK['month'] / 12)
      fleetUK['month_cos'] = np.cos(2 * np.pi * fleetUK['month'] / 12)
      fleetUK.drop('month', axis=1, inplace=True)
      
    • 检查One-Hot编码后的特征基数:如果某个分类特征的类别过多,One-Hot会导致维度爆炸,可尝试目标编码(Target Encoding)替代,避免维度冗余。
  • 特征重要性分析:用XGB的内置方法查看特征重要性,确认是否有特征完全无贡献:
    import xgboost as xgb
    xgb.plot_importance(my_model_1)
    plt.show()
    
    如果大部分特征重要性接近0,说明特征无法有效预测目标,需要重新挖掘交叉特征(比如year*month、gear_category*species等)。
  • 数据分布验证:查看训练集和测试集的目标变量、特征分布是否一致,避免分层抽样不足导致的分布偏移:
    # 对比训练/测试集Y的分布
    plt.hist(y_train, bins=30, alpha=0.5, label='Train')
    plt.hist(y_test, bins=30, alpha=0.5, label='Test')
    plt.legend()
    plt.show()
    
    如果分布差异大,改用分层抽样(将Y分箱后,用train_test_split的stratify参数)。

3. XGBoost的针对性调参

如果基准模型效果正常,XGB效果差,重点调以下参数:

  • 学习率与树数量:降低learning_rate(比如0.01),增加n_estimators(比如1000),配合早停:
    my_model_2 = XGBRegressor(
        random_state=22,
        learning_rate=0.01,
        n_estimators=1000,
        early_stopping_rounds=50
    )
    my_model_2.fit(X_train, y_train, eval_set=[(X_test, y_test)], verbose=False)
    
  • 树的复杂度:调整max_depth(默认6,可尝试3-10)、min_child_weight(控制叶子节点样本数,避免过拟合)、subsample/colsample_bytree(随机采样样本/特征,增加模型鲁棒性)。
  • 目标函数:XGB回归默认用reg:squarederror,如果Y是对数转换后的,可尝试reg:gamma或reg:tweedie,适配偏态分布。

4. 数据本身的可能性

如果以上调整都无效,考虑:

  • 数据集是否存在因果性缺失:比如预测value_per_ton的核心特征(如市场价格、捕捞成本)未包含在数据中,导致模型无法学习到有效规律。
  • 数据是否存在标签噪声:检查清洗后的Y值是否有异常,比如漏检的不合理极端值,或者标签本身标注错误。

内容的提问来源于stack exchange,提问作者RachZ

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 15:19:32