优化Numpy矩阵扫描方法:高效计算关联规则指标
优化商品关联规则计算的Numpy实现
问题背景
我有一个表示交易与商品的二进制二维Numpy数组(行代表交易,列代表商品),需要针对每个商品完成以下计算:
- 购买该商品的交易总数
- 与该商品一同被购买的所有商品及对应交易
- 基于上述数据计算商品间的支持度(support)、置信度(confidence)和提升度(lift)
目前用嵌套循环实现,但大数据集下耗时过长,想知道在必须完整扫描矩阵的场景下,怎么实现最优的扫描方式。
原实现代码
import numpy as np results = [] data = np.array([[0,1,0,1],[1,1,0,1],[0,0,0,1],[1,1,1,1]]) N = data.shape[0] #total transactions for i in range(4): # for each item i #transactions% where item i was purchased mask_i = data[:,i] == 1 support_i = data[mask_i].shape[0]/N for j in range(4): # check all items j #transactions where both i and j are purchased mask_j = data[:,j] == 1 support_j = data[mask_j].shape[0]/N #transactions% where j is purchased mask_c = np.logical_and(mask_i,mask_j) # mask for where both i and j were purchased confidence = (data[mask_c].shape[0]/N) / support_i lift = confidence/support_j results.append([i,j,confidence,lift]) #store combination
原代码输出结果
[[0, 0, 1.0, 2.0], [0, 1, 1.0, 1.3333333333333333], [0, 2, 0.5, 2.0], [0, 3, 1.0, 1.0], [1, 0, 0.6666666666666666, 1.3333333333333333], [1, 1, 1.0, 1.3333333333333333], [1, 2, 0.3333333333333333, 1.3333333333333333], [1, 3, 1.0, 1.0], [2, 0, 1.0, 2.0], [2, 1, 1.0, 1.3333333333333333], [2, 2, 1.0, 4.0], [2, 3, 1.0, 1.0], [3, 0, 0.5, 1.0], [3, 1, 0.75, 1.0], [3, 2, 0.25, 1.0], [3, 3, 1.0, 1.0]]
优化实现方案
核心思路是利用Numpy的向量化运算替代嵌套循环,只需要对原始矩阵做几次全局计算,就能批量得出所有商品组合的指标,彻底避免逐元素循环的开销。
步骤说明
预计算全局统计量
- 单个商品的购买交易数:直接按列求和,
item_counts = data.sum(axis=0) - 单个商品的支持度:购买交易数除以总交易数,
support = item_counts / N - 商品两两共现交易数:通过矩阵转置相乘得到,
co_occur = data.T @ data,其中co_occur[i][j]代表同时购买商品i和j的交易数
- 单个商品的购买交易数:直接按列求和,
批量计算指标
- 置信度:共现交易数除以商品i的购买交易数,利用广播机制批量计算所有组合:
confidence = co_occur / item_counts[:, None] - 提升度:置信度除以商品j的支持度,同样通过广播完成:
lift = confidence / support
- 置信度:共现交易数除以商品i的购买交易数,利用广播机制批量计算所有组合:
完整优化代码
import numpy as np data = np.array([[0,1,0,1],[1,1,0,1],[0,0,0,1],[1,1,1,1]]) N = data.shape[0] # 预计算全局统计量 item_counts = data.sum(axis=0) # 每个商品的购买交易数 support = item_counts / N # 每个商品的支持度 co_occur = data.T @ data # 两两商品的共现交易数矩阵 # 批量计算置信度和提升度 confidence = co_occur / item_counts[:, None] # 形状(M, M),M为商品数量 lift = confidence / support # 广播运算生成所有组合的提升度 # 整理为与原格式一致的结果列表 results = [] for i in range(len(support)): for j in range(len(support)): results.append([i, j, confidence[i,j], lift[i,j]]) # 输出结果 print(results)
优化效果
- 时间复杂度从原实现的O(M²N)降至O(MN + M²)(M为商品数,N为交易数),大数据集下性能提升显著
- 完全利用Numpy底层的C语言优化运算,避免Python循环的解释器开销
- 代码逻辑更简洁,可读性和维护性更强
内容的提问来源于stack exchange,提问作者rhamzaali
相关产品推荐
相关产品推荐

