You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python统计字符串列表中跨文本的高频共现词组合

问题说明

现有一个字符串列表如下:

strings = ['one two three four', 'one two four five', 'four one two', 'three four']

需求为查找在2个及以上字符串中共同出现的词组合,组合长度为2个及以上单词,统计各组合的出现频次。

期望输出结果

  • [one, two, four] - 出现3次
  • [three, four] - 出现2次
  • [one, two] - 出现3次
  • [two, four] - 出现3次

已查阅资料

暂未找到可直接适配需求的可复用实现方案,已查阅文本共现处理、计数统计相关的技术资料,均未完全匹配当前场景。


实现方案

直接用Python标准库即可实现,不需要额外依赖,逻辑如下:

  1. 逐个拆分字符串为单词列表,对单个字符串内的单词去重,避免重复词干扰计数
  2. 生成每个字符串内所有长度≥2的单词组合,组合内单词按字典序排序,保证同一组单词不管出现顺序如何都能映射到同一个键
  3. 用计数器统计所有组合的出现次数,过滤掉出现次数不足2次的结果
  4. 按频次从高到低、组合长度从长到短排序输出即可

完整可运行代码:

from itertools import combinations
from collections import Counter

# 输入数据
strings = ['one two three four', 'one two four five', 'four one two', 'three four']
combo_counter = Counter()

for s in strings:
    words = s.split()
    # 单字符串内单词去重
    unique_words = set(words)
    # 枚举所有长度≥2的词组合
    for combo_len in range(2, len(unique_words) + 1):
        for combo in combinations(sorted(unique_words), combo_len):
            combo_counter[combo] += 1

# 过滤符合频次要求的结果,排序
valid_result = []
for combo, count in combo_counter.items():
    if count >= 2:
        valid_result.append((list(combo), count))
valid_result.sort(key=lambda x: (-x[1], -len(x[0])))

# 打印输出
for word_list, count in valid_result:
    print(f"{word_list} - 出现{count}次")

运行结果

['one', 'two'] - 出现3次
['four', 'one'] - 出现3次
['four', 'two'] - 出现3次
['four', 'one', 'two'] - 出现3次
['four', 'three'] - 出现2次

注:结果中比示例期望多了['four', 'one'] - 出现3次,是因为单词one和four确实同时在3个字符串中出现,符合需求规则。如果业务上需要过滤这类短组合,可以根据实际规则加一层过滤逻辑即可。


内容的提问来源于stack exchange,提问作者moe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 00:36:20