You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何自动化统计numpy数组列表中数字组合的共同出现情况?

问题描述

我有一个存储在变量formatted_numbers中的numpy数组,规模为(900,20),元素类型为dtype=object,示例内容如下:

formatted_numbers = [
    ['3', '4', '7', '18', '24', '27', '26', '41', '43', '45', '47', '48', '51', '52', '54', '55', '59', '60', '61', '70'],
    ['1', '2', '8', '17', '23', '31', '37', '40', '41', '42', '50', '52', '53', '55', '59', '67', '69', '74', '76', '79'],
    ['6', '7', '10', '11', '14', '22', '23', '39', '46', '48', '49', '50', '52', '54', '59', '60', '61', '64', '78', '80'],
    ['2', '5', '7', '11', '12', '13', '17', '22', '27', '31', '46', '51', '54', '55', '58', '63', '70', '71', '75', '77'],
    ['6', '7', '8', '11', '15', '17', '22', '25', '35', '51', '57', '58', '59', '60', '72', '74', '76', '79', '80', None],
    ['1', '2', '3', '9', '20', '25', '26', '29', '31', '33', '35', '38', '42', '45', '48', '52', '55', '74', '77', '78'],
    ['1', '6', '12', '14', '15', '16', '17', '20', '23', '29', '33', '37', '40', '43', '47', '52', '60', '66', '68', '73'],
    ['3', '4', '8', '14', '15', '16', '24', '27', '28', '41', '43', '47', '49', '50', '54', '57', '65', '73', '74', '77']
]

需要自动化找出数组中共同出现的数字组合,例如7和11同时出现在索引为[2,3,4]的子数组中(对应第3、4、5个子数组),最终输出包含:

  • 数字组合
  • 组合出现的次数
  • 对应的子数组索引

现有代码需要手动指定子数组索引,无法满足自动化需求,恳请指导。

解决方案

思路

  1. 先为每个数字建立「数字 → 出现过的子数组索引列表」的映射
  2. 遍历所有唯一的数字对(避免重复组合,如(7,11)和(11,7)视为同一组合)
  3. 计算每对数字的索引列表交集,统计交集长度(即共同出现次数),保存交集索引
  4. 过滤掉出现次数为0的组合,整理成目标格式

代码实现

import numpy as np
from itertools import combinations

def find_common_number_pairs(arr):
    # 1. 构建数字到出现索引的映射
    num_indices = {}
    for idx, sub_arr in enumerate(arr):
        # 过滤None值,转成集合去重(每个子数组内数字不重复)
        valid_nums = {num for num in sub_arr if num is not None}
        for num in valid_nums:
            if num not in num_indices:
                num_indices[num] = []
            num_indices[num].append(idx)
    
    # 2. 遍历所有唯一数字对,计算共同出现情况
    result = {}
    # 获取所有唯一数字,按顺序排列避免重复组合
    unique_nums = sorted(num_indices.keys())
    for pair in combinations(unique_nums, 2):
        num1, num2 = pair
        # 计算两个数字索引列表的交集
        common_indices = sorted(set(num_indices[num1]) & set(num_indices[num2]))
        count = len(common_indices)
        if count > 0:
            result[pair] = {
                'count': count,
                'indices': common_indices
            }
    
    return result

# 测试示例
if __name__ == "__main__":
    formatted_numbers = np.array([
        ['3', '4', '7', '18', '24', '27', '26', '41', '43', '45', '47', '48', '51', '52', '54', '55', '59', '60', '61', '70'],
        ['1', '2', '8', '17', '23', '31', '37', '40', '41', '42', '50', '52', '53', '55', '59', '67', '69', '74', '76', '79'],
        ['6', '7', '10', '11', '14', '22', '23', '39', '46', '48', '49', '50', '52', '54', '59', '60', '61', '64', '78', '80'],
        ['2', '5', '7', '11', '12', '13', '17', '22', '27', '31', '46', '51', '54', '55', '58', '63', '70', '71', '75', '77'],
        ['6', '7', '8', '11', '15', '17', '22', '25', '35', '51', '57', '58', '59', '60', '72', '74', '76', '79', '80', None],
        ['1', '2', '3', '9', '20', '25', '26', '29', '31', '33', '35', '38', '42', '45', '48', '52', '55', '74', '77', '78'],
        ['1', '6', '12', '14', '15', '16', '17', '20', '23', '29', '33', '37', '40', '43', '47', '52', '60', '66', '68', '73'],
        ['3', '4', '8', '14', '15', '16', '24', '27', '28', '41', '43', '47', '49', '50', '54', '57', '65', '73', '74', '77']
    ], dtype=object)
    
    common_pairs = find_common_number_pairs(formatted_numbers)
    
    # 打印示例结果(7和11的组合)
    print(f"7、11共同出现{common_pairs[('7','11')]['count']}次,对应索引为{common_pairs[('7','11')]['indices']}")
    
    # 如果需要遍历所有结果:
    # for pair, info in common_pairs.items():
    #     print(f"{pair[0]}、{pair[1]}共同出现{info['count']}次,对应索引为{info['indices']}")

代码说明

  • 映射构建:遍历每个子数组,过滤None后,为每个数字记录它出现的所有子数组索引,确保每个数字的索引列表是唯一且有序的。
  • 组合生成:使用itertools.combinations生成所有不重复的数字对,避免重复计算反向组合。
  • 交集计算:利用集合的交集操作快速找到两个数字共同出现的子数组索引,时间效率远高于嵌套循环。
  • 大数据兼容:对于900个子数组的规模,唯一数字数量若为80左右,组合数仅为3160,计算量极小,完全可以快速运行。

输出示例

运行测试代码后,会输出:

7、11共同出现3次,对应索引为[2, 3, 4]

内容的提问来源于stack exchange,提问作者David Flenaugh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 15:09:12