You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用CountVectorizer处理数字时,为何小于10的数字无法被识别?

问题:CountVectorizer处理数字列表时,小于10的数字无法被正常识别?

我有一个数字列表,想用CountVectorizer处理,代码如下:

from sklearn.feature_extraction.text import CountVectorizer

def x(n):
   return str(n)

sentences = [5,10,15,10,5,10]

vectorizer = CountVectorizer(preprocessor= x, analyzer="word")
vectorizer.fit(sentences)

vectorizer.vocabulary_

输出结果:

{'10': 0, '15': 1}

执行转换代码:

vectorizer.transform(sentences).toarray()

输出结果:

array([[0, 0],
   [1, 0],
   [0, 1],
   [1, 0],
   [0, 0],
   [1, 0]], dtype=int64)

可以看到小于10的数字(比如5)没有被正常处理,这是为什么?


原因分析

CountVectorizer有个默认参数token_pattern,它的默认值是r"(?u)\b\w\w+\b"。这个正则表达式要求匹配的token(词)至少包含2个字符。

你把数字5转成字符串后是"5",只有1个字符,不符合默认的token匹配规则,所以被过滤掉了,没有被加入词汇表,自然在转换时也不会被统计。

解决方法

修改token_pattern参数,让它能匹配单个字符的词,将参数设置为r"(?u)\b\w+\b"即可,这个正则表达式会匹配任意长度的字母数字组合(包括单个字符)。

修改后的代码:

from sklearn.feature_extraction.text import CountVectorizer

def x(n):
   return str(n)

sentences = [5,10,15,10,5,10]

# 修改token_pattern参数
vectorizer = CountVectorizer(preprocessor=x, analyzer="word", token_pattern=r"(?u)\b\w+\b")
vectorizer.fit(sentences)

print(vectorizer.vocabulary_)
# 输出:{'5': 0, '10': 1, '15': 2}

print(vectorizer.transform(sentences).toarray())
# 输出:
# array([[1, 0, 0],
#        [0, 1, 0],
#        [0, 0, 1],
#        [0, 1, 0],
#        [1, 0, 0],
#        [0, 1, 0]], dtype=int64)

这样所有数字都能被正确识别和统计了。

内容的提问来源于stack exchange,提问作者saraafr

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 07:07:22