You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按美国郡分组运行Decision Tree Regression并验证准确率提升效果

按郡分组训练决策树回归模型实现方案

核心逻辑

用pandas的groupby方法按county字段拆分数据,对每个分组单独执行数据集拆分、模型训练、效果评估的全流程,最后汇总所有郡的结果和全局训练结果做对比即可。

实现代码

import pandas as pd
import numpy as np
from sklearn.tree import DecisionTreeRegressor
from sklearn.model_selection import train_test_split

# 样例数据生成(和你提供的一致)
df3 = pd.DataFrame({
  'y': np.random.randn(20),
  'a': np.random.randn(20), 
  'b': np.random.randn(20),
  'color': ['alf', 'bet', 'sar', 'tep'] * 5,
  'county': ['a', 'b'] * 10
})

# 定义每个郡的处理函数
def process_county(group_df, min_sample=10):
    # 过滤样本量过小的郡,避免拆分报错
    if len(group_df) < min_sample:
        return {
            'county': group_df['county'].iloc[0],
            'sample_count': len(group_df),
            'r2_score': np.nan,
            'model': None
        }
    # 拆分特征和标签,分组后county是常量,直接删除即可
    X = group_df.drop(['y', 'county'], axis=1)
    # 分类特征做独热编码处理,可根据自己的字段调整
    X = pd.get_dummies(X, drop_first=True)
    y = group_df['y']
    # 拆分训练测试集
    X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=0)
    # 训练DTR模型
    regressor = DecisionTreeRegressor(
        max_depth=10, 
        max_features='auto', 
        min_samples_leaf=2, # 样例数据量少,调小参数避免报错,自有数据集可改回原参数5
        min_samples_split=2, 
        random_state=42
    )
    regressor.fit(X_train, y_train)
    score = regressor.score(X_test, y_test)
    return {
        'county': group_df['county'].iloc[0],
        'sample_count': len(group_df),
        'r2_score': score,
        'model': regressor
    }

# 按county分组运行处理函数,汇总结果
result_list = []
for county_name, group in df3.groupby('county'):
    res = process_county(group)
    result_list.append(res)

# 转成DataFrame方便查看结果
result_df = pd.DataFrame(result_list)
print(result_df[['county', 'sample_count', 'r2_score']])

# 计算分组模型的平均R2,和全局模型对比
print("分组训练平均R2得分:", result_df['r2_score'].mean())

# 全局模型得分(用于对比)
X_global = pd.get_dummies(df3.drop(['y'], axis=1), drop_first=True)
y_global = df3['y']
X_train_g, X_test_g, y_train_g, y_test_g = train_test_split(X_global, y_global, test_size=0.2, random_state=0)
global_reg = DecisionTreeRegressor(max_depth=10, max_features='auto', min_samples_leaf=5, min_samples_split=5, random_state=42)
global_reg.fit(X_train_g, y_train_g)
print("全局训练R2得分:", global_reg.score(X_test_g, y_test_g))

注意事项

  • 分类特征(如样例中的color字段)不管是分组训练还是全局训练,都需要提前做独热编码或者标签编码处理
  • 如果部分郡的样本量太少,建议跳过或者和相邻郡合并,否则训练出来的模型泛化性很差,参考意义不大
  • 对比效果的时候不要只看平均R2,也要看不同郡的得分分布,判断是大部分郡效果都提升,还是只有部分郡有提升

内容的提问来源于stack exchange,提问作者bohontw

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 13:24:05