使用imblearn Pipeline报错:'Pipeline'对象无'_check_fit_params'属性
问题:imblearn Pipeline调用fit_resample时触发AttributeError
代码示例
from imblearn.over_sampling import SMOTE from imblearn.under_sampling import RandomUnderSampler from imblearn.pipeline import Pipeline # 定义特征与目标变量 X = df.drop('infected', axis=1) y = df['infected'] # 划分训练集与测试集 X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42) # 定义采样策略 over = SMOTE(sampling_strategy=0.5) # 将少数类过采样至多数类的50% under = RandomUnderSampler(sampling_strategy=0.8) # 将多数类欠采样至原规模的80% pipeline = Pipeline(steps=[('o', over), ('u', under)]) # 执行重采样 X_resampled, y_resampled = pipeline.fit_resample(X_train, y_train) # 输出新的类别分布 print("Resampled class distribution:", pd.Series(y_resampled).value_counts())
报错信息
AttributeError Traceback (most recent call last) Cell In[7], line 19 16 pipeline = Pipeline(steps=[('o', over), ('u', under)]) 18 # Apply the resampling ---> 19 X_resampled, y_resampled = pipeline.fit_resample(X_train, y_train) 21 # Show the new class distribution 22 print("Resampled class distribution:", pd.Series(y_resampled).value_counts()) File ~\anaconda3\Lib\site-packages\imblearn\pipeline.py:372, in Pipeline.fit_resample(self, X, y, **fit_params) 342 """Fit the model and sample with the final estimator. 343 344 Fits all the transformers/samplers one after the other and (...) 369 Transformed target. 370 """ 371 self._validate_params() ---> 372 fit_params_steps = self._check_fit_params(**fit_params) 373 Xt, yt = self._fit(X, y, **fit_params_steps) 374 last_step = self._final_estimator AttributeError: 'Pipeline' object has no attribute '_check_fit_params'
解决方案
这个错误核心原因是imbalanced-learn与scikit-learn版本不兼容,即使更新所有包也可能出现版本匹配错位的情况,以下是两种解决思路:
1. 强制安装版本匹配的包组合
imbalanced-learn的版本需要和scikit-learn严格对应,比如:
- imbalanced-learn 0.11.x ↔ scikit-learn 1.3.x
- imbalanced-learn 0.10.x ↔ scikit-learn 1.2.x
执行以下命令安装匹配版本(以1.3.x/0.11.x为例):
pip install -U imbalanced-learn==0.11.0 scikit-learn==1.3.0
或者让pip自动处理依赖匹配:
pip install -U "imbalanced-learn[all]"
2. 绕开imblearn Pipeline,手动依次执行采样
不用Pipeline,直接按顺序调用SMOTE和RandomUnderSampler的fit_resample方法:
# 先执行过采样 X_over, y_over = over.fit_resample(X_train, y_train) # 再执行欠采样 X_resampled, y_resampled = under.fit_resample(X_over, y_over) # 输出分布 print("Resampled class distribution:", pd.Series(y_resampled).value_counts())
内容的提问来源于stack exchange,提问作者Varutri Parihar
相关产品推荐
相关产品推荐

