You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

优化Numpy矩阵扫描方法:高效计算关联规则指标

优化商品关联规则计算的Numpy实现

问题背景

我有一个表示交易与商品的二进制二维Numpy数组(行代表交易,列代表商品),需要针对每个商品完成以下计算:

  • 购买该商品的交易总数
  • 与该商品一同被购买的所有商品及对应交易
  • 基于上述数据计算商品间的支持度(support)、置信度(confidence)和提升度(lift)

目前用嵌套循环实现,但大数据集下耗时过长,想知道在必须完整扫描矩阵的场景下,怎么实现最优的扫描方式。

原实现代码

import numpy as np
results = []
data = np.array([[0,1,0,1],[1,1,0,1],[0,0,0,1],[1,1,1,1]])
N = data.shape[0] #total transactions
for i in range(4): # for each item i
    #transactions% where item i was purchased
    mask_i = data[:,i] == 1
    support_i = data[mask_i].shape[0]/N
    for j in range(4): # check all items j
        #transactions where both i and j are purchased
        mask_j = data[:,j] == 1 
        support_j = data[mask_j].shape[0]/N #transactions% where j is purchased
        mask_c = np.logical_and(mask_i,mask_j) # mask for where both i and j were purchased
        confidence = (data[mask_c].shape[0]/N) / support_i
        lift = confidence/support_j
        results.append([i,j,confidence,lift]) #store combination

原代码输出结果

[[0, 0, 1.0, 2.0], [0, 1, 1.0, 1.3333333333333333], [0, 2, 0.5, 2.0], [0, 3, 1.0, 1.0], [1, 0, 0.6666666666666666, 1.3333333333333333], [1, 1, 1.0, 1.3333333333333333], [1, 2, 0.3333333333333333, 1.3333333333333333], [1, 3, 1.0, 1.0], [2, 0, 1.0, 2.0], [2, 1, 1.0, 1.3333333333333333], [2, 2, 1.0, 4.0], [2, 3, 1.0, 1.0], [3, 0, 0.5, 1.0], [3, 1, 0.75, 1.0], [3, 2, 0.25, 1.0], [3, 3, 1.0, 1.0]]

优化实现方案

核心思路是利用Numpy的向量化运算替代嵌套循环,只需要对原始矩阵做几次全局计算,就能批量得出所有商品组合的指标,彻底避免逐元素循环的开销。

步骤说明

  1. 预计算全局统计量

    • 单个商品的购买交易数:直接按列求和,item_counts = data.sum(axis=0)
    • 单个商品的支持度:购买交易数除以总交易数,support = item_counts / N
    • 商品两两共现交易数:通过矩阵转置相乘得到,co_occur = data.T @ data,其中co_occur[i][j]代表同时购买商品i和j的交易数
  2. 批量计算指标

    • 置信度:共现交易数除以商品i的购买交易数,利用广播机制批量计算所有组合:confidence = co_occur / item_counts[:, None]
    • 提升度:置信度除以商品j的支持度,同样通过广播完成:lift = confidence / support

完整优化代码

import numpy as np

data = np.array([[0,1,0,1],[1,1,0,1],[0,0,0,1],[1,1,1,1]])
N = data.shape[0]

# 预计算全局统计量
item_counts = data.sum(axis=0)  # 每个商品的购买交易数
support = item_counts / N       # 每个商品的支持度
co_occur = data.T @ data        # 两两商品的共现交易数矩阵

# 批量计算置信度和提升度
confidence = co_occur / item_counts[:, None]  # 形状(M, M),M为商品数量
lift = confidence / support                   # 广播运算生成所有组合的提升度

# 整理为与原格式一致的结果列表
results = []
for i in range(len(support)):
    for j in range(len(support)):
        results.append([i, j, confidence[i,j], lift[i,j]])

# 输出结果
print(results)

优化效果

  • 时间复杂度从原实现的O(M²N)降至O(MN + M²)(M为商品数,N为交易数),大数据集下性能提升显著
  • 完全利用Numpy底层的C语言优化运算,避免Python循环的解释器开销
  • 代码逻辑更简洁,可读性和维护性更强

内容的提问来源于stack exchange,提问作者rhamzaali

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 20:32:42