从边网络元组列表生成含最高频值(支持并列)的字典实现方案
问题:提取边网络列表中出现频率最高的节点及计数(支持并列)
需求:将包含数千个元组的边网络列表(示例:[('2975', '6384'), ('2975', '530'), ('7443', '1107983'), ('3534', '530')])转换为字典,返回出现频率最高的元素及其计数,且支持并列情况(预期结果:{'2975': 2, '530': 2})。
原实现代码
from collections import Counter nodes = [('2975', '6384'), ('2975', '530'), ('7443', '1107983'), ('3534', '530')] highest_nodes = {} # converting nodes data = Counter(list(sum(nodes, ()))).most_common() val = data[0][1] # get the value of n-1th item for a, b in list(takewhile(lambda x: x[1] >= val, data)): highest_nodes.setdefault(a, []).append(b) # This is returning values as a list containing the item, need to extract the int from it return highest_nodes
问题分析
原代码中使用highest_nodes.setdefault(a, []).append(b)会将每个节点的计数存入列表,导致结果格式不符合预期(比如{'2975': [2], '530': [2]}),同时缺少itertools.takewhile的导入语句,运行会报错。
修正方案
- 导入
itertools.takewhile模块; - 直接对字典键赋值单个计数值,而非追加到列表:
from collections import Counter from itertools import takewhile nodes = [('2975', '6384'), ('2975', '530'), ('7443', '1107983'), ('3534', '530')] highest_nodes = {} # 展开元组并统计节点频率 data = Counter(sum(nodes, ())).most_common() if not data: # 处理空列表的边界情况 return highest_nodes max_count = data[0][1] # 筛选出所有频率等于最高值的节点 for node, count in takewhile(lambda x: x[1] == max_count, data): highest_nodes[node] = count print(highest_nodes) # 输出: {'2975': 2, '530': 2}
说明
sum(nodes, ())可以直接展开所有元组为一个扁平元组,无需转成列表;- 用
x[1] == max_count替代>=更严谨,因为most_common()返回的结果是按频率降序排列的,后续元素的频率只会小于等于第一个元素; - 增加了空列表的边界处理,避免索引越界错误。
内容的提问来源于stack exchange,提问作者Mars
相关产品推荐
相关产品推荐

