基于AIC的Python逐步回归及statsmodels对应R step()函数咨询
Hey there! Let's tackle your two questions about AIC-based stepwise regression in Python—this is a common task when you want to trim down a regression model to the most impactful features.
There are two main approaches here: rolling your own implementation (which gives you full control) or using third-party libraries that simplify the process. Let's start with a manual backward stepwise example, since it's intuitive and uses statsmodels which you mentioned:
手动实现向后逐步回归(基于AIC)
Backward stepwise starts with the full model (all features) and iteratively removes the feature that leads to the lowest AIC, stopping when removing any remaining feature would increase the AIC.
Here's a code snippet to do this:
import statsmodels.api as sm import pandas as pd # 假设我们有一个数据集,y是目标变量,X是特征矩阵 data = pd.read_csv("your_data.csv") y = data["target"] X = data.drop("target", axis=1) X = sm.add_constant(X) # 添加截距项 # 拟合全模型 full_model = sm.OLS(y, X).fit() current_features = X.columns.tolist() best_aic = full_model.aic while True: aic_values = [] # 遍历每个特征,尝试移除它并计算AIC for feature in current_features: if feature == "const": continue # 不要移除截距项 temp_features = [f for f in current_features if f != feature] temp_model = sm.OLS(y, X[temp_features]).fit() aic_values.append((temp_model.aic, feature)) # 找到移除后AIC最小的特征 aic_values.sort() new_aic, removed_feature = aic_values[0] # 如果新AIC比当前最优AIC小,就移除该特征 if new_aic < best_aic: best_aic = new_aic current_features.remove(removed_feature) print(f"Removed feature: {removed_feature}, New AIC: {best_aic:.2f}") else: # 没有可以降低AIC的特征了,停止循环 break # 拟合最终模型 final_model = sm.OLS(y, X[current_features]).fit() print(final_model.summary())
If you want forward stepwise (starting with no features and adding the one that lowers AIC the most each time) or bidirectional stepwise (both adding and removing), you can adjust the loop logic similarly—just start with a minimal model and iterate through adding features instead.
step()函数的等效实现是什么?statsmodels.api中是否存在现成函数? Great question! R's step() function is super convenient because it handles forward, backward, and bidirectional stepwise out of the box with AIC as the default criterion.
关于statsmodels的现成函数
Unfortunately, statsmodels doesn't have a direct equivalent to R's step() built into the core library. The team has discussed adding this feature, but as of now, you'll need to implement it manually (like the example above) or use a third-party library.
第三方库替代方案
One popular option is mlxtend's SequentialFeatureSelector, which can be configured to use AIC as the evaluation metric. Here's how you can use it for backward stepwise:
from mlxtend.feature_selection import SequentialFeatureSelector as SFS from sklearn.linear_model import LinearRegression import statsmodels.api as sm import pandas as pd data = pd.read_csv("your_data.csv") y = data["target"] X = data.drop("target", axis=1) # 我们用sklearn的LinearRegression作为基础模型,然后用AIC来评分 def aic_scorer(estimator, X, y): # 转换为statsmodels的OLS模型来计算AIC X_const = sm.add_constant(X) model = sm.OLS(y, X_const).fit() return -model.aic # 因为SFS默认最大化评分,所以取负AIC # 初始化SFS,设置为向后选择,选择最优特征数 sfs = SFS( LinearRegression(), k_features="best", forward=False, # 向后选择 floating=False, # 纯向后,不是双向 scoring=aic_scorer, cv=0 # 不需要交叉验证,因为我们用AIC ) sfs = sfs.fit(X, y) print("Selected features:", sfs.k_feature_names_) # 拟合最终模型 final_X = sm.add_constant(X[list(sfs.k_feature_names_)]) final_model = sm.OLS(y, final_X).fit() print(final_model.summary())
This gets you pretty close to R's step() behavior. If you want bidirectional stepwise, just set floating=True in the SFS parameters.
内容的提问来源于stack exchange,提问作者Chenying Gao

