You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Linear SVC Pipeline特征提取异常:文本分类高频特征获取失败

文本分类高频特征提取问题

我在理解文本分类的高频特征时遇到了困难,试了两种方法:

第一种方法代码

def print_top10(vectorizer, clf, class_labels):
    """Prints features with the highest coefficient values, per class"""
    feature_names = vectorizer.get_feature_names_out()
    for i, class_label in enumerate(class_labels):
        top10 = np.argsort(clf.coef_[i])[-10:]
        print("%s: %s" % (class_label,
              " ".join(feature_names[j] for j in top10)))
class_labels=clf.classes_

第二种方法代码

def printNMostInformative(vectorizer, clf, N):
    feature_names = vectorizer.get_feature_names()
    coefs_with_fns = sorted(zip(clf.coef_[0], feature_names))
    topClass1 = coefs_with_fns[:N]
    topClass2 = coefs_with_fns[:-(N + 1):-1]
    print("Class 1 best: ")
    for feat in topClass1:
        print(feat)
    print("Class 2 best: ")
    for feat in topClass2:
        print(feat)

但两种方法只返回了准确率和空特征列表:

accuracy: 0.37922705314009664
Top 10 features used to predict: 
Class 1 best: 
(0.008202041988712563, '')
Class 2 best: 
(0.008202041988712563, '')

我的完整代码基于一篇机器学习spaCy相关的Notebook修改,仅替换了新数据集。


内容的提问来源于stack exchange,提问作者Gabry

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 17:24:48