You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用CountVectorizer在Python中基于自定义词汇表计算词出现频次的问询

使用CountVectorizer基于自定义词汇表计算文档词频

嘿,我来帮你搞定用CountVectorizer统计自定义多词词汇出现次数的事儿!默认的CountVectorizer只会拆单个词,没法直接识别多词短语,所以得稍微调整一下,下面是完整的解决方案:

核心思路

默认的CountVectorizer会把文本拆分成单个词汇,没法直接匹配你自定义的多词短语,所以我们需要自定义分词器来识别这些短语,同时指定vocabulary参数来限定统计范围。

完整代码实现

1. 导入依赖并准备数据

from sklearn.feature_extraction.text import CountVectorizer
import re

# 你的样本文档
docs = [
    'And that was the fallacy. Once I was free to talk with staff members',
    'In the new, stripped-down, every-job-counts business climate, these human',
    'Another reality makes emotional intelligence ever more crucial',
    'The globalization of the workforce puts a particular premium on emotional',
    'As business changes, so do the traits needed to excel. Data tracking'
]

# 自定义词汇表(支持多词短语)
my_vocabulary = ['was the fallacy', 'free to', 'stripped-down', 'emotional intelligence', 'globalization of the workforce']

2. 自定义分词器

这个分词器会优先匹配长短语(避免短短语拆分长短语),同时处理标点和大小写:

def custom_tokenizer(text):
    # 移除文本中的标点符号(可根据需求调整,如果你需要保留词汇中的标点就去掉这步)
    text_clean = re.sub(r'[^\w\s-]', '', text)
    # 按短语的词数倒序排列,优先匹配更长的短语
    sorted_vocab = sorted(my_vocabulary, key=lambda x: len(x.split()), reverse=True)
    # 构建正则匹配规则,匹配所有自定义词汇
    pattern = re.compile(r'\b(' + '|'.join(re.escape(vocab) for vocab in sorted_vocab) + r')\b', re.IGNORECASE)
    # 提取所有匹配的短语
    tokens = pattern.findall(text_clean)
    # 统一转成小写(保证大小写不影响统计)
    return [token.lower() for token in tokens]

3. 初始化CountVectorizer并计算词频

# 初始化向量器:指定自定义词汇表和分词器,关闭默认的token_pattern
vectorizer = CountVectorizer(vocabulary=my_vocabulary, tokenizer=custom_tokenizer, token_pattern=None)

# 拟合文档并生成词频矩阵
word_counts = vectorizer.fit_transform(docs)

# 输出结果
print("自定义词汇表:", vectorizer.get_feature_names_out())
print("文档词频矩阵:\n", word_counts.toarray())

输出结果解释

运行代码后你会得到这样的输出:

自定义词汇表: ['was the fallacy' 'free to' 'stripped-down' 'emotional intelligence' 'globalization of the workforce']
文档词频矩阵:
 [[1 1 0 0 0]
 [0 0 1 0 0]
 [0 0 0 1 0]
 [0 0 0 0 1]
 [0 0 0 0 0]]
  • 每一行对应一个文档,每一列对应自定义词汇表中的一个短语
  • 数字1表示该短语在对应文档中出现1次,0表示未出现

关键注意事项

  • 短语优先级排序:一定要把长短语放在前面匹配,比如如果词汇表同时有free和free to,先匹配free to才不会把它拆成free和to
  • 标点处理:如果你的自定义词汇包含标点(比如stripped-down),re.escape()会自动转义特殊字符,不用担心匹配问题
  • 大小写兼容:分词器中统一转小写,确保Free To和free to会被视为同一个短语

内容的提问来源于stack exchange,提问作者nightrain

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 08:43:07