You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为字典键按出现次数反向分配概率?概率求和异常修复

为字典标签分配反向概率的代码修正方案

需求说明

需要为字典的键(标签)分配概率,规则为:出现次数越低的键获得越高的概率,出现次数越高的键获得越低的概率。目标字典如下:

my_dict = {0: 21, 1: 36, 2: 13, 3: 344, 4: 171, 5: 10, 6: 7, 7: 24, 8: 15, 9: 14, 10: 77, 11: 7, 12: 434, 13: 6, 14: 38, 15: 328, 16: 149, 17: 12, 18: 67, 19: 85, 20: 33, 21: 19, 22: 13, 23: 477, 24: 9, 25: 206, 26: 226, 27: 48, 28: 135, 29: 42, 30: 273, 31: 11, 32: 61, 33: 11, 34: 378, 35: 32, 36: 10, 37: 237, 38: 248, 39: 64, 40: 7, 41: 74, 42: 17, 43: 30, 44: 12, 45: 44, 46: 197, 47: 314, 48: 118, 49: 40, 50: 89, 51: 6, 52: 260, 53: 18, 54: 5, 55: 5, 56: 5, 57: 455, 58: 25, 59: 23, 60: 70, 61: 179, 62: 98, 63: 9, 64: 163, 65: 102, 66: 8, 67: 188, 68: 5, 69: 500, 70: 8, 71: 142, 72: 216, 73: 6, 74: 299, 75: 286, 76: 56, 77: 156, 78: 123, 79: 58, 80: 27, 81: 20, 82: 93, 83: 29, 84: 361, 85: 26, 86: 15, 87: 396, 88: 112, 89: 415, 90: 46, 91: 53, 92: 16, 93: 6, 94: 81, 95: 22, 96: 129, 97: 51, 98: 35, 99: 107}

原代码问题分析

你编写的代码存在两个核心问题:

  1. 变量名错误:使用了未定义的dictionary,应该替换为my_dict。
  2. 归一化逻辑错误:用所有标签的出现次数总和作为分母,但分子是基于排序位次的权重值,两者量级不匹配,导致概率总和无法为1。

原代码:

sorted_dict = {key: value for key, value in sorted(dictionary.items(), key=lambda item: item[1])}

# Step 2: Calculate the probability of each key based on the sorted order
total_count = sum(sorted_dict.values())
probabilities = {key: (len(sorted_dict) - index)/total_count for index, (key, value) in enumerate(sorted_dict.items())}

# Print the probabilities

 for key, probability in probabilities.items():
     print(f"Key: {key}, Probability: {probability}")
    
print(sum(probabilities.values()))

修正方案

要实现概率总和为1,需基于权重的总和进行归一化,而非原出现次数的总和。具体步骤:

  1. 按标签出现次数升序排序,得到有序的键值对列表。
  2. 为每个标签分配权重:排序越靠前(次数越少),权重越高,权重公式为总标签数 - 当前索引。
  3. 计算所有权重的总和,用每个标签的权重除以总权重,得到归一化后的概率。

修正后的完整代码

# 1. 按出现次数升序排序,得到有序列表(Python3.7+字典有序,但用列表更兼容)
sorted_items = sorted(my_dict.items(), key=lambda item: item[1])
total_labels = len(sorted_items)

# 2. 计算每个标签的权重及总权重
total_weight = total_labels * (total_labels + 1) // 2  # 等差数列求和公式,等价于sum(total_labels - idx for idx in range(total_labels))

# 3. 生成概率字典
probabilities = {key: (total_labels - idx) / total_weight for idx, (key, val) in enumerate(sorted_items)}

# 4. 输出结果
for key, probability in probabilities.items():
    print(f"Key: {key}, Probability: {probability:.6f}")

# 验证概率总和(应为1)
print(f"\nProbability sum: {sum(probabilities.values()):.6f}")

代码说明

  • 排序后的sorted_items确保次数最少的标签排在最前面,获得最大的权重(total_labels)。
  • 总权重使用等差数列求和公式计算,效率更高,以此为分母归一化后,概率总和必然为1。
  • 使用.6f格式化输出,让概率值更易读。

内容的提问来源于stack exchange,提问作者Siddeshwar Raghavan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 14:10:35