You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何让CountVectorizer()忽略大小写停用词且不转换文本大小写?

解决CountVectorizer保留大小写同时过滤所有形式停用词的问题

默认情况下,CountVectorizer使用的英文停用词表全是小写形式,当你设置lowercase=False时,内置的停用词过滤只会匹配小写的停用词(比如the),无法识别The、THE这类变体。要实现保留原文本大小写且过滤所有大小写形式停用词的需求,你可以通过自定义tokenizer来解决:

具体实现步骤

  1. 获取Sklearn默认的英文停用词集合
  2. 自定义tokenizer函数,拆分文本后检查每个token的小写形式是否在停用词表中,仅保留非停用词的原大小写token
  3. 将自定义tokenizer传入CountVectorizer,同时关闭内置的停用词过滤

完整代码示例

from sklearn.feature_extraction.text import CountVectorizer
from sklearn.feature_extraction.text import _check_stop_list

# 获取默认的英文停用词集合
stop_words = _check_stop_list("english")

def custom_tokenizer(text):
    # 使用CountVectorizer默认的分词逻辑拆分文本
    default_tokenizer = CountVectorizer().build_tokenizer()
    tokens = default_tokenizer(text)
    # 过滤掉小写形式属于停用词的token,保留原大小写
    return [token for token in tokens if token.lower() not in stop_words]

# 初始化CountVectorizer
ngram_range = (1, 1)  # 可根据你的需求调整
vectorizer = CountVectorizer(
    tokenizer=custom_tokenizer,
    lowercase=False,
    ngram_range=ngram_range,
    stop_words=None  # 关闭内置停用词过滤,避免重复处理
)

测试验证

比如输入文本:

"The cat is chasing THE mouse with the dog"

经过处理后,生成的有效token会是:["cat", "chasing", "mouse", "dog"],所有大小写形式的the都被过滤,同时其他词保留了原有的大小写格式。

如果需要处理ngram中的停用词(比如(2,2)的ngram),这个逻辑同样适用——因为tokenizer先过滤了单个停用词,后续生成的ngram不会包含以停用词为组成部分的短语。

内容的提问来源于stack exchange,提问作者suprita shankar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 03:33:09