使用CountVectorizer在Python中基于自定义词汇表计算词出现频次的问询
使用CountVectorizer基于自定义词汇表计算文档词频
嘿,我来帮你搞定用CountVectorizer统计自定义多词词汇出现次数的事儿!默认的CountVectorizer只会拆单个词,没法直接识别多词短语,所以得稍微调整一下,下面是完整的解决方案:
核心思路
默认的CountVectorizer会把文本拆分成单个词汇,没法直接匹配你自定义的多词短语,所以我们需要自定义分词器来识别这些短语,同时指定vocabulary参数来限定统计范围。
完整代码实现
1. 导入依赖并准备数据
from sklearn.feature_extraction.text import CountVectorizer import re # 你的样本文档 docs = [ 'And that was the fallacy. Once I was free to talk with staff members', 'In the new, stripped-down, every-job-counts business climate, these human', 'Another reality makes emotional intelligence ever more crucial', 'The globalization of the workforce puts a particular premium on emotional', 'As business changes, so do the traits needed to excel. Data tracking' ] # 自定义词汇表(支持多词短语) my_vocabulary = ['was the fallacy', 'free to', 'stripped-down', 'emotional intelligence', 'globalization of the workforce']
2. 自定义分词器
这个分词器会优先匹配长短语(避免短短语拆分长短语),同时处理标点和大小写:
def custom_tokenizer(text): # 移除文本中的标点符号(可根据需求调整,如果你需要保留词汇中的标点就去掉这步) text_clean = re.sub(r'[^\w\s-]', '', text) # 按短语的词数倒序排列,优先匹配更长的短语 sorted_vocab = sorted(my_vocabulary, key=lambda x: len(x.split()), reverse=True) # 构建正则匹配规则,匹配所有自定义词汇 pattern = re.compile(r'\b(' + '|'.join(re.escape(vocab) for vocab in sorted_vocab) + r')\b', re.IGNORECASE) # 提取所有匹配的短语 tokens = pattern.findall(text_clean) # 统一转成小写(保证大小写不影响统计) return [token.lower() for token in tokens]
3. 初始化CountVectorizer并计算词频
# 初始化向量器:指定自定义词汇表和分词器,关闭默认的token_pattern vectorizer = CountVectorizer(vocabulary=my_vocabulary, tokenizer=custom_tokenizer, token_pattern=None) # 拟合文档并生成词频矩阵 word_counts = vectorizer.fit_transform(docs) # 输出结果 print("自定义词汇表:", vectorizer.get_feature_names_out()) print("文档词频矩阵:\n", word_counts.toarray())
输出结果解释
运行代码后你会得到这样的输出:
自定义词汇表: ['was the fallacy' 'free to' 'stripped-down' 'emotional intelligence' 'globalization of the workforce'] 文档词频矩阵: [[1 1 0 0 0] [0 0 1 0 0] [0 0 0 1 0] [0 0 0 0 1] [0 0 0 0 0]]
- 每一行对应一个文档,每一列对应自定义词汇表中的一个短语
- 数字
1表示该短语在对应文档中出现1次,0表示未出现
关键注意事项
- 短语优先级排序:一定要把长短语放在前面匹配,比如如果词汇表同时有
free和free to,先匹配free to才不会把它拆成free和to - 标点处理:如果你的自定义词汇包含标点(比如
stripped-down),re.escape()会自动转义特殊字符,不用担心匹配问题 - 大小写兼容:分词器中统一转小写,确保
Free To和free to会被视为同一个短语
内容的提问来源于stack exchange,提问作者nightrain
相关产品推荐
相关产品推荐

