如何为字典键按出现次数反向分配概率?概率求和异常修复
为字典标签分配反向概率的代码修正方案
需求说明
需要为字典的键(标签)分配概率,规则为:出现次数越低的键获得越高的概率,出现次数越高的键获得越低的概率。目标字典如下:
my_dict = {0: 21, 1: 36, 2: 13, 3: 344, 4: 171, 5: 10, 6: 7, 7: 24, 8: 15, 9: 14, 10: 77, 11: 7, 12: 434, 13: 6, 14: 38, 15: 328, 16: 149, 17: 12, 18: 67, 19: 85, 20: 33, 21: 19, 22: 13, 23: 477, 24: 9, 25: 206, 26: 226, 27: 48, 28: 135, 29: 42, 30: 273, 31: 11, 32: 61, 33: 11, 34: 378, 35: 32, 36: 10, 37: 237, 38: 248, 39: 64, 40: 7, 41: 74, 42: 17, 43: 30, 44: 12, 45: 44, 46: 197, 47: 314, 48: 118, 49: 40, 50: 89, 51: 6, 52: 260, 53: 18, 54: 5, 55: 5, 56: 5, 57: 455, 58: 25, 59: 23, 60: 70, 61: 179, 62: 98, 63: 9, 64: 163, 65: 102, 66: 8, 67: 188, 68: 5, 69: 500, 70: 8, 71: 142, 72: 216, 73: 6, 74: 299, 75: 286, 76: 56, 77: 156, 78: 123, 79: 58, 80: 27, 81: 20, 82: 93, 83: 29, 84: 361, 85: 26, 86: 15, 87: 396, 88: 112, 89: 415, 90: 46, 91: 53, 92: 16, 93: 6, 94: 81, 95: 22, 96: 129, 97: 51, 98: 35, 99: 107}
原代码问题分析
你编写的代码存在两个核心问题:
- 变量名错误:使用了未定义的
dictionary,应该替换为my_dict。 - 归一化逻辑错误:用所有标签的出现次数总和作为分母,但分子是基于排序位次的权重值,两者量级不匹配,导致概率总和无法为1。
原代码:
sorted_dict = {key: value for key, value in sorted(dictionary.items(), key=lambda item: item[1])} # Step 2: Calculate the probability of each key based on the sorted order total_count = sum(sorted_dict.values()) probabilities = {key: (len(sorted_dict) - index)/total_count for index, (key, value) in enumerate(sorted_dict.items())} # Print the probabilities for key, probability in probabilities.items(): print(f"Key: {key}, Probability: {probability}") print(sum(probabilities.values()))
修正方案
要实现概率总和为1,需基于权重的总和进行归一化,而非原出现次数的总和。具体步骤:
- 按标签出现次数升序排序,得到有序的键值对列表。
- 为每个标签分配权重:排序越靠前(次数越少),权重越高,权重公式为
总标签数 - 当前索引。 - 计算所有权重的总和,用每个标签的权重除以总权重,得到归一化后的概率。
修正后的完整代码
# 1. 按出现次数升序排序,得到有序列表(Python3.7+字典有序,但用列表更兼容) sorted_items = sorted(my_dict.items(), key=lambda item: item[1]) total_labels = len(sorted_items) # 2. 计算每个标签的权重及总权重 total_weight = total_labels * (total_labels + 1) // 2 # 等差数列求和公式,等价于sum(total_labels - idx for idx in range(total_labels)) # 3. 生成概率字典 probabilities = {key: (total_labels - idx) / total_weight for idx, (key, val) in enumerate(sorted_items)} # 4. 输出结果 for key, probability in probabilities.items(): print(f"Key: {key}, Probability: {probability:.6f}") # 验证概率总和(应为1) print(f"\nProbability sum: {sum(probabilities.values()):.6f}")
代码说明
- 排序后的
sorted_items确保次数最少的标签排在最前面,获得最大的权重(total_labels)。 - 总权重使用等差数列求和公式计算,效率更高,以此为分母归一化后,概率总和必然为1。
- 使用
.6f格式化输出,让概率值更易读。
内容的提问来源于stack exchange,提问作者Siddeshwar Raghavan
相关产品推荐
相关产品推荐

