使用CountVectorizer处理数字时,为何小于10的数字无法被识别?
问题:CountVectorizer处理数字列表时,小于10的数字无法被正常识别?
我有一个数字列表,想用CountVectorizer处理,代码如下:
from sklearn.feature_extraction.text import CountVectorizer def x(n): return str(n) sentences = [5,10,15,10,5,10] vectorizer = CountVectorizer(preprocessor= x, analyzer="word") vectorizer.fit(sentences) vectorizer.vocabulary_
输出结果:
{'10': 0, '15': 1}
执行转换代码:
vectorizer.transform(sentences).toarray()
输出结果:
array([[0, 0], [1, 0], [0, 1], [1, 0], [0, 0], [1, 0]], dtype=int64)
可以看到小于10的数字(比如5)没有被正常处理,这是为什么?
原因分析
CountVectorizer有个默认参数token_pattern,它的默认值是r"(?u)\b\w\w+\b"。这个正则表达式要求匹配的token(词)至少包含2个字符。
你把数字5转成字符串后是"5",只有1个字符,不符合默认的token匹配规则,所以被过滤掉了,没有被加入词汇表,自然在转换时也不会被统计。
解决方法
修改token_pattern参数,让它能匹配单个字符的词,将参数设置为r"(?u)\b\w+\b"即可,这个正则表达式会匹配任意长度的字母数字组合(包括单个字符)。
修改后的代码:
from sklearn.feature_extraction.text import CountVectorizer def x(n): return str(n) sentences = [5,10,15,10,5,10] # 修改token_pattern参数 vectorizer = CountVectorizer(preprocessor=x, analyzer="word", token_pattern=r"(?u)\b\w+\b") vectorizer.fit(sentences) print(vectorizer.vocabulary_) # 输出:{'5': 0, '10': 1, '15': 2} print(vectorizer.transform(sentences).toarray()) # 输出: # array([[1, 0, 0], # [0, 1, 0], # [0, 0, 1], # [0, 1, 0], # [1, 0, 0], # [0, 1, 0]], dtype=int64)
这样所有数字都能被正确识别和统计了。
内容的提问来源于stack exchange,提问作者saraafr
相关产品推荐
相关产品推荐

