Python statsmodels是否有支持direction='both'的R stepAIC等价实现
解答
statsmodels 官方没有提供和R语言MASS包中stepAIC(model, direction="both")完全等价的开箱即用工具,目前公开可查的第三方实现大多仅支持前向或后向的单向逐步选择。下面给出可直接复用的双向逐步AIC选择实现,完全兼容statsmodels的广义线性模型接口,支持你正在使用的sm.families.NegativeBinomial()负二项分布家族,逻辑完全对齐R中direction="both"的搜索规则。
双向搜索核心规则
双向选择和单向选择的核心差异是每轮迭代同时评估两类变量调整的收益:
- 对当前模型已经纳入的所有自变量,逐个测试移除该变量后新模型的AIC值
- 对全量候选变量中尚未纳入模型的变量,逐个测试加入该变量后新模型的AIC值
- 从所有可能的单步调整(加一个变量/删一个变量)中,选择能让AIC下降幅度最大的操作更新模型
- 重复上述过程,直到任何加、删变量的操作都无法让AIC继续降低时终止搜索。该逻辑天然支持此前被移除的变量在后续步骤重新入模、此前纳入的变量在后续步骤被移除,和R的
stepAIC双向模式行为完全一致。
可直接运行的实现代码
import statsmodels.api as sm import pandas as pd import numpy as np def step_aic_both(initial_model, exog_full, endog, family, trace=True): """ 兼容statsmodels GLM的双向逐步AIC选择,对齐R MASS包stepAIC direction="both"逻辑 参数: initial_model: 初始拟合完成的statsmodels GLM模型对象 exog_full: 包含全量候选自变量的DataFrame,若需要常数项请提前通过sm.add_constant添加 endog: 模型因变量,支持序列、数组格式 family: statsmodels支持的分布族对象,例如sm.families.NegativeBinomial()、sm.families.Binomial()等 trace: 是否打印每一步的变量调整过程与AIC变化 返回: 筛选完成的最优拟合GLM模型 """ # 初始化当前模型状态 current_used_cols = list(initial_model.params.index) current_min_aic = initial_model.aic all_candidate_cols = list(exog_full.columns) while True: possible_moves = [] # 枚举所有删除变量的可能操作 for col in current_used_cols: # 默认保留常数项不参与筛选,不需要可删除该判断 if col == "const": continue test_cols = [c for c in current_used_cols if c != col] test_model = sm.GLM(endog, exog_full[test_cols], family=family).fit() possible_moves.append( (test_model.aic, "remove", col, test_cols, test_model) ) # 枚举所有新增变量的可能操作 unused_cols = [c for c in all_candidate_cols if c not in current_used_cols] for col in unused_cols: test_cols = current_used_cols + [col] test_model = sm.GLM(endog, exog_full[test_cols], family=family).fit() possible_moves.append( (test_model.aic, "add", col, test_cols, test_model) ) # 按AIC从小到大排序所有可能操作 possible_moves.sort(key=lambda x: x[0]) best_move_aic, best_move_type, best_move_col, best_move_cols, best_move_model = possible_moves[0] # 没有能让AIC下降的操作,终止搜索 if best_move_aic >= current_min_aic: if trace: print(f"搜索结束,最终模型AIC = {current_min_aic:.4f},入模变量列表:{current_used_cols}") break # 执行最优操作,更新当前模型状态 current_min_aic = best_move_aic current_used_cols = best_move_cols if trace: print(f"完成操作:{best_move_type} 变量 `{best_move_col}`,当前AIC = {current_min_aic:.4f},入模变量:{current_used_cols}") # 拟合并返回最终最优模型 final_model = sm.GLM(endog, exog_full[current_used_cols], family=family).fit() return final_model
使用示例
# 负二项回归双向逐步选择示例 # 1. 准备数据:假设df为你的分析数据集,target是因变量,候选自变量为x1~x6 # y = df["target"] # X_candidate = sm.add_constant(df[["x1", "x2", "x3", "x4", "x5", "x6"]]) # 2. 定义初始模型,可根据需求选择空模型(仅常数项)或全变量模型作为起点 # # 以仅含常数项的空模型为起点示例: # initial_model = sm.GLM(y, X_candidate[["const"]], family=sm.families.NegativeBinomial()).fit() # 3. 运行双向逐步选择 # best_nb_model = step_aic_both( # initial_model=initial_model, # exog_full=X_candidate, # endog=y, # family=sm.families.NegativeBinomial() # ) # 4. 查看最终模型结果 # print(best_nb_model.summary())
使用说明
- 代码默认不将常数项
const纳入筛选范围,始终保留在模型中,如果你的场景不需要常数项,直接删除if col == "const": continue这段判断即可 - 初始模型的选择和R中
stepAIC逻辑一致:既可以从仅含常数项的空模型开始搜索,也可以从全变量模型、或你提前指定的任意变量组合的初始模型开始搜索 - 所有模型的AIC值完全调用statsmodels模型内置的
.aic属性计算,不存在自定义计算偏差,对所有statsmodels支持的GLM分布族通用,不限于负二项分布。
内容的提问来源于stack exchange,提问作者FLX
相关产品推荐
相关产品推荐

