如何在NLP多分类任务的Sklearn Pipeline中正确使用SMOTE
问题原因分析
- 第一个报错原因:你在Pipeline前直接调用SMOTE时,输入的X还是原始文本字符串,SMOTE是基于特征向量空间的距离生成样本,要求输入必须是数值型特征,自然会报无法转float的错误。
- 第二个报错原因:sklearn原生的Pipeline要求所有中间步骤必须同时实现
fit和transform方法,而SMOTE属于采样类组件,只有fit_resample方法,没有符合要求的transform接口,所以无法直接加入sklearn原生Pipeline。
解决方法
用imbalanced-learn库自带的Pipeline替代sklearn原生Pipeline,它专门适配了采样类组件,会自动在训练阶段执行采样逻辑,预测阶段自动跳过采样步骤,避免测试集数据泄露。
完整修改方案
首先确保已安装依赖:
pip install imbalanced-learn
修改你的代码如下:
# 新增导入 from imblearn.pipeline import Pipeline as ImbPipeline from imblearn.over_sampling import SMOTE # 原有文本处理、数据集拆分代码保持不变 X = df['product_description'] y = df['class'] X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.3, random_state=42 ) def text_process(mess): STOPWORDS = stopwords.words("english") nopunc = [char for char in mess if char not in string.punctuation] nopunc = "".join(nopunc) return " ".join([word for word in nopunc.split() if word.lower() not in STOPWORDS]) # 替换为imblearn的Pipeline,新增SMOTE步骤 pipe = ImbPipeline( steps=[ ("vect", CountVectorizer(analyzer= text_process)), ("feature_selection", SelectKBest(chi2, k=20)), ("smote", SMOTE(random_state=42)), # SMOTE放在数值特征生成之后、分类器之前 ("polynomial", PolynomialFeatures(2)), # 新增class_weight和max_iter参数,优化不平衡场景下的收敛和效果 ("reg", LogisticRegression(class_weight="balanced", max_iter=1000)), ] ) pipe.fit(X_train, y_train) y_pred = pipe.predict(X_test) print(classification_report(y_test, y_pred))
补充说明
- SMOTE步骤必须放在所有文本转数值的步骤之后,此时输入已经是数值型特征矩阵,符合SMOTE的输入要求。
- imblearn的Pipeline会自动保证:仅在
fit阶段对训练集做过采样,predict阶段不会对测试集做任何采样操作,完全避免数据泄露问题。 - 如果对采样比例有自定义需求,可以调整SMOTE的
sampling_strategy参数,比如传入字典指定每个类别的采样后样本量。
内容的提问来源于stack exchange,提问作者dekio
相关产品推荐
相关产品推荐

