如何统计字典中相似列表组合的出现次数并生成命名列表?
问题:统计列表组合的出现次数并生成对应命名列表
输入的字典集合:
{'HH1': ['x'], 'HH2': ['y', 'x'], 'HH3': ['x', 'z'], 'HH4': ['x'], 'HH5': ['x'], 'HH6': ['x'], 'HH7': ['x'], 'HH8': ['x', 'y', 'z'], 'HH9': ['x'], 'HH10': ['x', 'y'], 'HH11': ['x'], 'HH12': ['x'], 'HH13': ['x'], 'HH14': ['x'], 'HH15': ['x', 'y'], 'HH16': ['x', 'y'], 'HH17': ['x', 'y'], 'HH18': ['x']}
需求
- 统计每个完全相同的元素组合的出现次数i(元素顺序不影响,比如
['x','y']和['y','x']视为同一组合) - 生成每个组合对应的结果,命名格式为
n=i
期望输出
n=11: ('x',) n=5: ('x', 'y') n=1: ('x', 'z') n=1: ('x', 'y', 'z')
两次尝试的问题
尝试1:遗漏组合
代码:
from collections import Counter from itertools import combinations # Counting occurrences of each combination of fuel types combination_count = Counter() unique_combinations = set() for fuel_list in hfuels.values(): for r in range(1, len(fuel_list) + 1): for combination in combinations(fuel_list, r): unique_combinations.add(tuple(sorted(combination))) # Creating new lists for each combination renamed_lists = {} for combination in unique_combinations: count = sum(1 for fuel_list in hfuels.values() if set(combination) == set(fuel_list)) if count: renamed_lists[f"n={count}"] = list(combination) # Printing the renamed lists for name, fuel_list in renamed_lists.items(): print(f"{name}: {fuel_list}")
执行结果:
n=1: ['z', 'y', 'x'] n=5: ['y', 'x'] n=11: ['x']
问题原因:
- 错误生成了所有子组合(比如从
['x','y','z']中拆分出('x','y')等子组合),而非统计原列表的完整组合。 - 用字典存储结果时,相同次数的组合会被覆盖(比如
('x','z')和('x','y','z')的次数都是1,后者会覆盖前者,导致前者丢失)。
尝试2:统计子组合次数而非完整组合次数
代码:
from collections import Counter from itertools import combinations # Counting occurrences of each combination of fuel types combination_count = Counter() unique_combinations = set() for fuel_list in hfuels.values(): for r in range(1, len(fuel_list) + 1): for combination in combinations(fuel_list, r): combination_count[tuple(sorted(combination))] += 1 unique_combinations.add(tuple(sorted(combination))) # Creating new lists for each combination renamed_lists = {} for combination in unique_combinations: count = combination_count[combination] if count: renamed_lists[f"n={count}"] = list(combination) # Printing the renamed lists for name, fuel_list in renamed_lists.items(): print(f"{name}: {fuel_list}")
执行结果:
n=1: ['z', 'y', 'x'] n=2: ['z'] n=6: ['y'] n=18: ['x']
问题原因:
代码统计的是所有子组合的出现次数(比如每个包含x的列表都会贡献一次('x')),而非原列表完整组合的出现次数,完全偏离需求。
正确实现方法
核心思路:直接将每个原列表转换为排序后的元组(解决列表不可哈希、元素顺序干扰的问题),然后用Counter统计每个元组的出现次数即可。
代码:
from collections import Counter hfuels = {'HH1': ['x'], 'HH2': ['y', 'x'], 'HH3': ['x', 'z'], 'HH4': ['x'], 'HH5': ['x'], 'HH6': ['x'], 'HH7': ['x'], 'HH8': ['x', 'y', 'z'], 'HH9': ['x'], 'HH10': ['x', 'y'], 'HH11': ['x'], 'HH12': ['x'], 'HH13': ['x'], 'HH14': ['x'], 'HH15': ['x', 'y'], 'HH16': ['x', 'y'], 'HH17': ['x', 'y'], 'HH18': ['x']} # 将每个列表转为排序后的元组,确保顺序不影响组合判断 sorted_combinations = [tuple(sorted(lst)) for lst in hfuels.values()] # 统计每个组合的出现次数 counts = Counter(sorted_combinations) # 按需求格式输出 for combo, count in counts.items(): print(f"n={count}: {combo}")
执行结果:
n=11: ('x',) n=5: ('x', 'y') n=1: ('x', 'z') n=1: ('x', 'y', 'z')
内容的提问来源于stack exchange,提问作者Nike
相关产品推荐
相关产品推荐

