You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用改进版CountVectorizer实现序列词频统计及单元格映射求和

实现方案

核心逻辑分为两步:

  • 先遍历全量文本序列,统计每个单词的全局总出现频次,构造词到频次的映射字典
  • 再遍历单个文本,将文本内的每个单词替换为对应全局频次,按需求拼接为目标格式或直接求和

方法1:纯Python用Counter实现(无需额外依赖)

from collections import Counter

# 输入文本序列
texts = ["dog cat mouse", " cat mouse", "mouse mouse cat"]

# 统计全局词频
all_words = []
for text in texts:
    all_words.extend(text.strip().split())
global_freq = Counter(all_words)

# 生成目标格式结果
result = []
for text in texts:
    words = text.strip().split()
    freq_str = "+".join([str(global_freq[word]) for word in words])
    result.append(freq_str)

print(result)
# 输出:['1+3+4', '3+4', '4+4+3']

如果需要直接得到求和后的数值,将拼接步骤替换为求和即可:

sum_val = sum([global_freq[word] for word in words])
# 对应输出为 [8, 7, 11]

方法2:基于CountVectorizer实现

from sklearn.feature_extraction.text import CountVectorizer
import numpy as np

texts = ["dog cat mouse", " cat mouse", "mouse mouse cat"]

# 拟合全量文本,统计全局词频
vec = CountVectorizer()
vec.fit(texts)
# 按列求和得到每个词的全局总出现次数
global_count = np.asarray(vec.transform(texts).sum(axis=0)).flatten()
# 构造词到频次的映射
word_to_freq = {word: global_count[idx] for word, idx in vec.vocabulary_.items()}

# 生成结果
result = []
for text in texts:
    words = text.strip().split()
    freq_str = "+".join([str(word_to_freq[word]) for word in words])
    result.append(freq_str)

print(result)
# 输出:['1+3+4', '3+4', '4+4+3']

之前用Counter未达到预期的原因通常是仅统计了单句内的词频,没有先完成全量语料的全局频次统计,按照上述流程先做全局统计再映射替换即可解决问题。

内容的提问来源于stack exchange,提问作者Yash Sing

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.07 10:18:03