如何修改Sklearn TfidfVectorizer的token_pattern以包含数字及符号token
Hey there! I see exactly what you're trying to do — the default behavior of TfidfVectorizer ignores those numeric tokens with symbols because its default token_pattern only targets letter-based words. Let's fix that with a simple regex adjustment.
核心问题原因
The default token_pattern is r'(?u)\b\w\w+\b', which only matches tokens that have at least two word characters (letters/underscores). That's why your numeric tokens like 1/8 or 1-1/4 are getting left out.
解决方案:调整token_pattern参数
We just need to update the regex pattern to include / and - as valid characters in tokens. Here's the pattern you should use: r'(?u)\b[\w/-]+\b'. Let's break this down:
(?u): Turns on Unicode-aware matching, so it works with any language's characters if needed\b: Ensures we capture full tokens (not partial fragments)[\w/-]+: Matches one or more of: word characters (letters, numbers, underscores),/, or-
修改后的完整代码
import pandas as pd from sklearn.feature_extraction.text import TfidfVectorizer data1 = ['1/8 wire','4 tube','1-1/4 brush'] dataset = pd.DataFrame(data1, columns=['des']) # 配置TfidfVectorizer,使用自定义token_pattern vectorizer1 = TfidfVectorizer(lowercase=False, token_pattern=r'(?u)\b[\w/-]+\b') tf_idf_matrix = pd.DataFrame( vectorizer1.fit_transform(dataset['des']).toarray(), columns=vectorizer1.get_feature_names_out() # 新版本sklearn推荐用这个方法,避免警告 ) # 打印词汇表验证结果 print(vectorizer1.get_feature_names_out())
预期输出
['1-1/4' '1/8' '4' 'brush' 'tube' 'wire']
小提示
If you're using an older version of scikit-learn (before 1.0), you can replace get_feature_names_out() with get_feature_names() — both will work, but the former is the updated, recommended method.
内容的提问来源于stack exchange,提问作者Ranjana Girish

