You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在NLP多分类任务的Sklearn Pipeline中正确使用SMOTE

问题原因分析
  • 第一个报错原因:你在Pipeline前直接调用SMOTE时,输入的X还是原始文本字符串,SMOTE是基于特征向量空间的距离生成样本,要求输入必须是数值型特征,自然会报无法转float的错误。
  • 第二个报错原因:sklearn原生的Pipeline要求所有中间步骤必须同时实现fit和transform方法,而SMOTE属于采样类组件,只有fit_resample方法,没有符合要求的transform接口,所以无法直接加入sklearn原生Pipeline。
解决方法

用imbalanced-learn库自带的Pipeline替代sklearn原生Pipeline,它专门适配了采样类组件,会自动在训练阶段执行采样逻辑,预测阶段自动跳过采样步骤,避免测试集数据泄露。

完整修改方案

首先确保已安装依赖:

pip install imbalanced-learn

修改你的代码如下:

# 新增导入
from imblearn.pipeline import Pipeline as ImbPipeline
from imblearn.over_sampling import SMOTE

# 原有文本处理、数据集拆分代码保持不变
X = df['product_description']
y = df['class']

X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.3, random_state=42
)

def text_process(mess):
    STOPWORDS = stopwords.words("english")
    nopunc = [char for char in mess if char not in string.punctuation]
    nopunc = "".join(nopunc)
    return " ".join([word for word in nopunc.split() if word.lower() not in STOPWORDS])

# 替换为imblearn的Pipeline,新增SMOTE步骤
pipe = ImbPipeline(
steps=[
    ("vect", CountVectorizer(analyzer= text_process)),
    ("feature_selection", SelectKBest(chi2, k=20)),
    ("smote", SMOTE(random_state=42)), # SMOTE放在数值特征生成之后、分类器之前
    ("polynomial", PolynomialFeatures(2)),
    # 新增class_weight和max_iter参数,优化不平衡场景下的收敛和效果
    ("reg", LogisticRegression(class_weight="balanced", max_iter=1000)),
]
)

pipe.fit(X_train, y_train)
y_pred = pipe.predict(X_test)
print(classification_report(y_test, y_pred))
补充说明
  • SMOTE步骤必须放在所有文本转数值的步骤之后,此时输入已经是数值型特征矩阵,符合SMOTE的输入要求。
  • imblearn的Pipeline会自动保证:仅在fit阶段对训练集做过采样,predict阶段不会对测试集做任何采样操作,完全避免数据泄露问题。
  • 如果对采样比例有自定义需求,可以调整SMOTE的sampling_strategy参数,比如传入字典指定每个类别的采样后样本量。

内容的提问来源于stack exchange,提问作者dekio

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 00:06:03