如何使用项矩阵扫描候选项集并统计其出现次数?
解决项矩阵中候选项集出现次数统计问题
问题说明
需要扫描布尔值组成的项矩阵(DataFrame格式),统计每个候选项集(集合类型)在矩阵中作为子集出现的次数——即每行中候选项集的所有元素对应列值均为True的总行数。由于普通集合不可哈希,无法直接作为字典键存储计数,需针对性处理。
解决方案一:用frozenset做字典键 + 逐候选集计算
把普通集合转为可哈希的frozenset作为字典键,再遍历每个候选项集计算出现次数:
import pandas as pd from mlxtend.preprocessing import TransactionEncoder # 初始化数据集与项矩阵 dataset = [['Milk', 'Onion', 'Nutmeg', 'Kidney Beans', 'Eggs', 'Yogurt'], ['Dill', 'Onion', 'Nutmeg', 'Kidney Beans', 'Eggs', 'Yogurt'], ['Milk', 'Apple', 'Kidney Beans', 'Eggs'], ['Milk', 'Unicorn', 'Corn', 'Kidney Beans', 'Yogurt'], ['Corn', 'Onion', 'Onion', 'Kidney Beans', 'Ice cream', 'Eggs']] te = TransactionEncoder() te_ary = te.fit(dataset).transform(dataset) df = pd.DataFrame(te_ary, columns=te.columns_) # 候选项集列表 candidate_set = [{'Apple', 'Milk'}, {'Eggs', 'Milk'}, {'Onion', 'Yogurt'}] # 转换为frozenset并统计次数 count_dict = {} for candidate in candidate_set: frozen_candidate = frozenset(candidate) # 计算每行中候选集所有元素均为True的行数 count = df[list(candidate)].all(axis=1).sum() count_dict[frozen_candidate] = count print(count_dict) # 输出:{frozenset({'Apple', 'Milk'}): 1, frozenset({'Eggs', 'Milk'}): 2, frozenset({'Onion', 'Yogurt'}): 2}
解决方案二:批量处理(更高效)
若候选项集数量较多,用列表推导式批量生成计数:
count_dict = { frozenset(candidate): df[list(candidate)].all(axis=1).sum() for candidate in candidate_set } print(count_dict)
关键逻辑说明
df[list(candidate)]:提取候选项集对应的矩阵列.all(axis=1):判断每行中所有列是否均为True(即候选项集是当前行的子集).sum():统计满足条件的总行数frozenset(candidate):将普通集合转为可哈希类型,作为字典键存储计数结果
内容的提问来源于stack exchange,提问作者Cooper
相关产品推荐
相关产品推荐

