如何自动化统计numpy数组列表中数字组合的共同出现情况?
问题描述
我有一个存储在变量formatted_numbers中的numpy数组,规模为(900,20),元素类型为dtype=object,示例内容如下:
formatted_numbers = [ ['3', '4', '7', '18', '24', '27', '26', '41', '43', '45', '47', '48', '51', '52', '54', '55', '59', '60', '61', '70'], ['1', '2', '8', '17', '23', '31', '37', '40', '41', '42', '50', '52', '53', '55', '59', '67', '69', '74', '76', '79'], ['6', '7', '10', '11', '14', '22', '23', '39', '46', '48', '49', '50', '52', '54', '59', '60', '61', '64', '78', '80'], ['2', '5', '7', '11', '12', '13', '17', '22', '27', '31', '46', '51', '54', '55', '58', '63', '70', '71', '75', '77'], ['6', '7', '8', '11', '15', '17', '22', '25', '35', '51', '57', '58', '59', '60', '72', '74', '76', '79', '80', None], ['1', '2', '3', '9', '20', '25', '26', '29', '31', '33', '35', '38', '42', '45', '48', '52', '55', '74', '77', '78'], ['1', '6', '12', '14', '15', '16', '17', '20', '23', '29', '33', '37', '40', '43', '47', '52', '60', '66', '68', '73'], ['3', '4', '8', '14', '15', '16', '24', '27', '28', '41', '43', '47', '49', '50', '54', '57', '65', '73', '74', '77'] ]
需要自动化找出数组中共同出现的数字组合,例如7和11同时出现在索引为[2,3,4]的子数组中(对应第3、4、5个子数组),最终输出包含:
- 数字组合
- 组合出现的次数
- 对应的子数组索引
现有代码需要手动指定子数组索引,无法满足自动化需求,恳请指导。
解决方案
思路
- 先为每个数字建立「数字 → 出现过的子数组索引列表」的映射
- 遍历所有唯一的数字对(避免重复组合,如
(7,11)和(11,7)视为同一组合) - 计算每对数字的索引列表交集,统计交集长度(即共同出现次数),保存交集索引
- 过滤掉出现次数为0的组合,整理成目标格式
代码实现
import numpy as np from itertools import combinations def find_common_number_pairs(arr): # 1. 构建数字到出现索引的映射 num_indices = {} for idx, sub_arr in enumerate(arr): # 过滤None值,转成集合去重(每个子数组内数字不重复) valid_nums = {num for num in sub_arr if num is not None} for num in valid_nums: if num not in num_indices: num_indices[num] = [] num_indices[num].append(idx) # 2. 遍历所有唯一数字对,计算共同出现情况 result = {} # 获取所有唯一数字,按顺序排列避免重复组合 unique_nums = sorted(num_indices.keys()) for pair in combinations(unique_nums, 2): num1, num2 = pair # 计算两个数字索引列表的交集 common_indices = sorted(set(num_indices[num1]) & set(num_indices[num2])) count = len(common_indices) if count > 0: result[pair] = { 'count': count, 'indices': common_indices } return result # 测试示例 if __name__ == "__main__": formatted_numbers = np.array([ ['3', '4', '7', '18', '24', '27', '26', '41', '43', '45', '47', '48', '51', '52', '54', '55', '59', '60', '61', '70'], ['1', '2', '8', '17', '23', '31', '37', '40', '41', '42', '50', '52', '53', '55', '59', '67', '69', '74', '76', '79'], ['6', '7', '10', '11', '14', '22', '23', '39', '46', '48', '49', '50', '52', '54', '59', '60', '61', '64', '78', '80'], ['2', '5', '7', '11', '12', '13', '17', '22', '27', '31', '46', '51', '54', '55', '58', '63', '70', '71', '75', '77'], ['6', '7', '8', '11', '15', '17', '22', '25', '35', '51', '57', '58', '59', '60', '72', '74', '76', '79', '80', None], ['1', '2', '3', '9', '20', '25', '26', '29', '31', '33', '35', '38', '42', '45', '48', '52', '55', '74', '77', '78'], ['1', '6', '12', '14', '15', '16', '17', '20', '23', '29', '33', '37', '40', '43', '47', '52', '60', '66', '68', '73'], ['3', '4', '8', '14', '15', '16', '24', '27', '28', '41', '43', '47', '49', '50', '54', '57', '65', '73', '74', '77'] ], dtype=object) common_pairs = find_common_number_pairs(formatted_numbers) # 打印示例结果(7和11的组合) print(f"7、11共同出现{common_pairs[('7','11')]['count']}次,对应索引为{common_pairs[('7','11')]['indices']}") # 如果需要遍历所有结果: # for pair, info in common_pairs.items(): # print(f"{pair[0]}、{pair[1]}共同出现{info['count']}次,对应索引为{info['indices']}")
代码说明
- 映射构建:遍历每个子数组,过滤
None后,为每个数字记录它出现的所有子数组索引,确保每个数字的索引列表是唯一且有序的。 - 组合生成:使用
itertools.combinations生成所有不重复的数字对,避免重复计算反向组合。 - 交集计算:利用集合的交集操作快速找到两个数字共同出现的子数组索引,时间效率远高于嵌套循环。
- 大数据兼容:对于900个子数组的规模,唯一数字数量若为80左右,组合数仅为3160,计算量极小,完全可以快速运行。
输出示例
运行测试代码后,会输出:
7、11共同出现3次,对应索引为[2, 3, 4]
内容的提问来源于stack exchange,提问作者David Flenaugh
相关产品推荐
相关产品推荐

