You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在sklearn CountVectorizer(char_wb模式)中移除含空格特征

解决CountVectorizer字符级n-gram含空格特征的问题

要去掉analyzer='char_wb'生成的含空格特征,有两种直接的解决办法:

方法一:自定义analyzer函数(推荐)

直接从源头生成仅包含单词内部字符的n-gram,避免产生带空格的边界特征。核心思路是先拆分文本为单个单词,再对每个单词提取指定长度的字符n-gram:

from sklearn.feature_extraction.text import CountVectorizer

def custom_char_analyzer(text):
    words = text.split()
    ngrams = []
    # 遍历指定的n-gram长度(4和5)
    for n in range(4, 6):
        for word in words:
            # 仅当单词长度大于等于n时生成n-gram
            if len(word) >= n:
                ngrams.extend([word[i:i+n] for i in range(len(word)-n+1)])
    return ngrams

# 使用自定义analyzer初始化向量器
vectorizer = CountVectorizer(binary=True, analyzer=custom_char_analyzer)
vectorizer.fit(['this is a plural'])
print(vectorizer.vocabulary_)

运行结果会得到不含空格的词汇表:

{'this': 0, 'plur': 1, 'lura': 2, 'ural': 3, 'plura': 4, 'lural': 5}

方法二:事后过滤词汇表

如果想保留原char_wb的逻辑但剔除含空格特征,可以在拟合后手动过滤词汇表:

from sklearn.feature_extraction.text import CountVectorizer

# 先按原配置拟合
vectorizer = CountVectorizer(binary=True, analyzer='char_wb', ngram_range=(4, 5))
vectorizer.fit(['this is a plural'])

# 过滤掉包含空格的特征
filtered_vocab = {ngram: idx for ngram, idx in vectorizer.vocabulary_.items() if ' ' not in ngram}

# 使用过滤后的词汇表创建新向量器
new_vectorizer = CountVectorizer(
    binary=True, 
    analyzer='char_wb', 
    ngram_range=(4, 5), 
    vocabulary=filtered_vocab
)
new_vectorizer.fit(['this is a plural'])
print(new_vectorizer.vocabulary_)

这个方法会得到和方法一相同的结果,但自定义analyzer更高效,因为它不会先生成那些不需要的带空格特征。

内容的提问来源于stack exchange,提问作者Ankit Bansal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 12:10:23