You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

从边网络元组列表生成含最高频值(支持并列)的字典实现方案

问题:提取边网络列表中出现频率最高的节点及计数(支持并列)

需求:将包含数千个元组的边网络列表(示例:[('2975', '6384'), ('2975', '530'), ('7443', '1107983'), ('3534', '530')])转换为字典,返回出现频率最高的元素及其计数,且支持并列情况(预期结果:{'2975': 2, '530': 2})。

原实现代码

from collections import Counter

nodes = [('2975', '6384'), ('2975', '530'), ('7443', '1107983'), ('3534', '530')]

highest_nodes = {}

# converting nodes
data = Counter(list(sum(nodes, ()))).most_common()
val = data[0][1]  # get the value of n-1th item
for a, b in list(takewhile(lambda x: x[1] >= val, data)):
    highest_nodes.setdefault(a, []).append(b)
# This is returning values as a list containing the item, need to extract the int from it
return highest_nodes

问题分析

原代码中使用highest_nodes.setdefault(a, []).append(b)会将每个节点的计数存入列表,导致结果格式不符合预期(比如{'2975': [2], '530': [2]}),同时缺少itertools.takewhile的导入语句,运行会报错。

修正方案

  1. 导入itertools.takewhile模块;
  2. 直接对字典键赋值单个计数值,而非追加到列表:
from collections import Counter
from itertools import takewhile

nodes = [('2975', '6384'), ('2975', '530'), ('7443', '1107983'), ('3534', '530')]

highest_nodes = {}

# 展开元组并统计节点频率
data = Counter(sum(nodes, ())).most_common()
if not data:  # 处理空列表的边界情况
    return highest_nodes
max_count = data[0][1]

# 筛选出所有频率等于最高值的节点
for node, count in takewhile(lambda x: x[1] == max_count, data):
    highest_nodes[node] = count

print(highest_nodes)  # 输出: {'2975': 2, '530': 2}

说明

  • sum(nodes, ())可以直接展开所有元组为一个扁平元组,无需转成列表;
  • 用x[1] == max_count替代>=更严谨,因为most_common()返回的结果是按频率降序排列的,后续元素的频率只会小于等于第一个元素;
  • 增加了空列表的边界处理,避免索引越界错误。

内容的提问来源于stack exchange,提问作者Mars

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 18:30:59