使用改进版CountVectorizer实现序列词频统计及单元格映射求和
实现方案
核心逻辑分为两步:
- 先遍历全量文本序列,统计每个单词的全局总出现频次,构造词到频次的映射字典
- 再遍历单个文本,将文本内的每个单词替换为对应全局频次,按需求拼接为目标格式或直接求和
方法1:纯Python用Counter实现(无需额外依赖)
from collections import Counter # 输入文本序列 texts = ["dog cat mouse", " cat mouse", "mouse mouse cat"] # 统计全局词频 all_words = [] for text in texts: all_words.extend(text.strip().split()) global_freq = Counter(all_words) # 生成目标格式结果 result = [] for text in texts: words = text.strip().split() freq_str = "+".join([str(global_freq[word]) for word in words]) result.append(freq_str) print(result) # 输出:['1+3+4', '3+4', '4+4+3']
如果需要直接得到求和后的数值,将拼接步骤替换为求和即可:
sum_val = sum([global_freq[word] for word in words]) # 对应输出为 [8, 7, 11]
方法2:基于CountVectorizer实现
from sklearn.feature_extraction.text import CountVectorizer import numpy as np texts = ["dog cat mouse", " cat mouse", "mouse mouse cat"] # 拟合全量文本,统计全局词频 vec = CountVectorizer() vec.fit(texts) # 按列求和得到每个词的全局总出现次数 global_count = np.asarray(vec.transform(texts).sum(axis=0)).flatten() # 构造词到频次的映射 word_to_freq = {word: global_count[idx] for word, idx in vec.vocabulary_.items()} # 生成结果 result = [] for text in texts: words = text.strip().split() freq_str = "+".join([str(word_to_freq[word]) for word in words]) result.append(freq_str) print(result) # 输出:['1+3+4', '3+4', '4+4+3']
之前用Counter未达到预期的原因通常是仅统计了单句内的词频,没有先完成全量语料的全局频次统计,按照上述流程先做全局统计再映射替换即可解决问题。
内容的提问来源于stack exchange,提问作者Yash Sing
相关产品推荐
相关产品推荐

