You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

CountVectorizer未遵循正则表达式规则,如何解决异常词条问题?

如何修复文档词项矩阵中的异常词条?

看起来你遇到的问题是生成的词项矩阵里出现了纯数字、数字组合甚至编码乱码的无效词条,比如20009、2001â这类。这些词条大多没有实际语义价值,我们可以从文本预处理、分词规则和词过滤三个层面来解决:

1. 过滤无字母的无效词条

你的StemmedCountVectorizer可以在词干提取后,直接过滤掉那些不包含任何字母的词——纯数字、乱码字符基本都符合这个特征。修改你的类实现:

from nltk.stem import SnowballStemmer
from sklearn.feature_extraction.text import CountVectorizer

stemmer = SnowballStemmer("english", ignore_stopwords=True)

class StemmedCountVectorizer(CountVectorizer):
    def build_analyzer(self):
        analyzer = super(StemmedCountVectorizer, self).build_analyzer()
        # 只保留包含至少一个字母的词
        return lambda doc: ([stemmer.stem(w) for w in analyzer(doc) 
                             if any(c.isalpha() for c in w)])

stemmed_count_vect = StemmedCountVectorizer(stop_words='english', 
                                            ngram_range=(1,1), 
                                            token_pattern=r'\b\w+\b', 
                                            min_df=1, 
                                            max_df=0.6)

2. 预处理清理编码乱码

像2001â这类是典型的Unicode编码错误,我们可以先对原始文本做规范化处理,把乱码字符转成可识别内容或直接去除:

import unicodedata

def clean_raw_text(text):
    # 规范化Unicode字符,忽略无法转成ASCII的乱码
    normalized_text = unicodedata.normalize('NFKD', text)
    cleaned_text = normalized_text.encode('ascii', 'ignore').decode('utf-8')
    return cleaned_text

# 先清理所有原始文本,再传入向量器处理
processed_corpus = [clean_raw_text(doc) for doc in your_original_corpus]
term_matrix = stemmed_count_vect.fit_transform(processed_corpus)

3. 优化分词规则从源头拦截

你当前的token_pattern=r'\b\w+\b'会匹配任何字母数字组合,我们可以修改规则,只匹配至少包含一个字母的词,从分词阶段就排除纯数字:

stemmed_count_vect = StemmedCountVectorizer(stop_words='english', 
                                            ngram_range=(1,1), 
                                            # 要求词以字母开头,后续可以跟字母或数字
                                            token_pattern=r'\b[a-zA-Z]+\w*\b', 
                                            min_df=1, 
                                            max_df=0.6)

组合方案建议

实际使用时,建议把这三个方法结合起来:先清理文本编码问题,再用优化的分词规则,最后在词干提取后做过滤,这样能最大程度减少无效词条的出现。

内容的提问来源于stack exchange,提问作者juliano.net

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:21:16