You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

关于开发Scikit-learn全回归器调用框架的技术咨询

解决Scikit-learn回归器框架的两个核心问题

嘿,这个需求我之前做模型对比的时候也碰到过,刚好能给你一套落地的方案!咱们一步步来:

1. 自动获取所有Scikit-learn回归器列表

Scikit-learn其实提供了现成的工具来枚举所有内置评估器,不用你手动一个个列出来。用sklearn.utils.all_estimators()就能搞定,再筛选出类型为回归器的即可:

from sklearn.utils import all_estimators
from sklearn.base import RegressorMixin

# 获取所有回归器类
all_regressors = all_estimators(type_filter="regressor")

# 整理成字典:键是模型名称,值是模型类
regressor_dict = {name: estimator for name, estimator in all_regressors if issubclass(estimator, RegressorMixin)}

# 打印看看有多少个回归器
print(f"总共找到 {len(regressor_dict)} 个回归器")

小提示:有些回归器可能需要额外依赖(比如IsotonicRegression),或是专门处理多输出任务的,你可以根据数据集情况,在筛选时加上额外条件(比如排除需要特殊参数初始化的模型)。

2. 批量运行模型、计算指标与超参数调优

拿到回归器列表后,接下来就是批量训练、评估、调参的流程,咱们拆成几个小模块:

2.1 定义评估指标计算函数

先写一个通用函数,用来计算你需要的RMSE、R-Sq和Adjusted R-Sq:

import numpy as np
from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_squared_error, r2_score

def evaluate_regressor(regressor_class, X, y):
    # 拆分数据集
    X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
    
    # 初始化模型(先默认参数,后面调参再优化)
    model = regressor_class()
    
    # 训练与预测
    model.fit(X_train, y_train)
    y_pred = model.predict(X_test)
    
    # 计算指标
    rmse = np.sqrt(mean_squared_error(y_test, y_pred))
    r2 = r2_score(y_test, y_pred)
    # 手动计算Adjusted R²(sklearn无直接实现)
    adjusted_r2 = 1 - (1 - r2) * (len(y_test) - 1) / (len(y_test) - X_test.shape[1] - 1)
    
    return {
        "model_name": regressor_class.__name__,
        "rmse": rmse,
        "r2": r2,
        "adjusted_r2": adjusted_r2
    }

2.2 批量运行所有回归器

遍历之前的回归器字典,逐个评估:

# 假设你的预处理后数据集是X和y
results = []
for name, regressor in regressor_dict.items():
    try:
        # 跳过需要嵌套其他模型的特殊回归器
        if name in ["MultiOutputRegressor", "StackingRegressor", "VotingRegressor"]:
            continue
        res = evaluate_regressor(regressor, X, y)
        results.append(res)
        print(f"完成评估:{name}")
    except Exception as e:
        print(f"评估 {name} 失败:{str(e)}")

# 转成DataFrame方便排序查看
import pandas as pd
results_df = pd.DataFrame(results).sort_values(by="rmse", ascending=True)
print(results_df)

2.3 超参数调优

对表现不错的模型,用GridSearchCV做调参优化(也可以用RandomizedSearchCV提高效率):

from sklearn.model_selection import GridSearchCV

# 定义模型与对应的参数网格(可按需扩展)
tune_configs = [
    {
        "model": regressor_dict["RandomForestRegressor"],
        "param_grid": {
            "n_estimators": [100, 200],
            "max_depth": [None, 10, 20]
        }
    },
    {
        "model": regressor_dict["GradientBoostingRegressor"],
        "param_grid": {
            "learning_rate": [0.01, 0.1],
            "n_estimators": [100, 200]
        }
    }
]

tuned_results = []
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
for config in tune_configs:
    model = config["model"]()
    grid_search = GridSearchCV(model, config["param_grid"], cv=5, scoring="neg_mean_squared_error")
    grid_search.fit(X_train, y_train)
    
    # 用最优模型重新评估
    best_model = grid_search.best_estimator_
    y_pred = best_model.predict(X_test)
    rmse = np.sqrt(mean_squared_error(y_test, y_pred))
    r2 = r2_score(y_test, y_pred)
    adjusted_r2 = 1 - (1 - r2) * (len(y_test) - 1) / (len(y_test) - X_test.shape[1] - 1)
    
    tuned_results.append({
        "model_name": model.__class__.__name__,
        "best_params": grid_search.best_params_,
        "tuned_rmse": rmse,
        "tuned_r2": r2,
        "tuned_adjusted_r2": adjusted_r2
    })

tuned_results_df = pd.DataFrame(tuned_results)
print(tuned_results_df)

这套流程下来,就基本实现了你想要的类似caret的模型对比、调参功能啦。记得所有模型要在相同预处理后的数据集上运行,这样对比结果才有效!

内容的提问来源于stack exchange,提问作者rishm.msc

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 08:29:01