You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于列表众数接近度与出现频率生成0-1区间评分?

问题描述

我需要为列表中的距离值生成0-1区间的评分,评分依据为该值与列表众数的接近度以及自身出现频率。目前尝试的Python代码无法准确体现其在列表中的价值:

def calc_racedistance_likeability_score(race_dist,list_of_distances):
    list_of_distances = [10, 5,5,5,5, 10, 16.09, 10, 10.7, 10,
                   10, 5, "Marathon", 15, 10, 10, 5, 10,
                   15, 10, 10,10, 5, 5, 5, 5, 5]

#calculate mode to get users most chosen race
    most_chosen_distance = statistics.mode(list_of_distances)    #get mode of list 

    if race_dist == most_chosen_distance:
        score = 1
    else:
        frequency_score = (list_of_distances.count(race_dist) / len(list_of_distances))
        accuracy_score = score = 1 - abs(race_dist - most_chosen_distance) / (race_dist + most_chosen_distance)
        score = (frequency_score * accuracy_score)

    print(f"score is {score}...for distance {race_dist}")

理想的评分需要满足:

  • 略微偏向更小数值(比如众数为10时,9的评分优于11)
  • 出现频率的权重高于与众数的接近度
  • 评分范围在0-1之间,越接近1越好

解决方案

核心调整思路

  1. 清理非数值数据:原列表中的"Marathon"会干扰众数计算,先将其转换为标准数值(42.195),过滤其他无效非数值项
  2. 偏向性接近度计算:对小于众数的距离额外加分,实现“偏小数值更优”的需求
  3. 加权权重分配:用加权求和替代相乘,给频率评分更高权重(比如6:4的比例),避免低频率直接拉低总分
  4. 归一化约束:确保最终评分严格落在0-1区间内

优化后的代码

import statistics

def calc_racedistance_likeability_score(race_dist, list_of_distances):
    # 预处理:统一转换数值,处理Marathon
    processed_distances = []
    for d in list_of_distances:
        if isinstance(d, str):
            if d.lower() == "marathon":
                processed_distances.append(42.195)
            continue
        processed_distances.append(float(d))
    
    # 计算众数与基准参数
    most_chosen = statistics.mode(processed_distances)
    total_count = len(processed_distances)
    max_dist = max(processed_distances)
    
    # 处理目标距离的类型转换
    if isinstance(race_dist, str):
        if race_dist.lower() == "marathon":
            race_dist = 42.195
        else:
            return 0.0  # 未知字符串返回0分
    
    # 众数直接得满分
    if race_dist == most_chosen:
        return 1.0
    
    # 计算频率评分:出现次数/总次数,天然在0-1区间
    freq_count = processed_distances.count(race_dist)
    frequency_score = freq_count / total_count
    
    # 计算带偏向的接近度评分
    diff = race_dist - most_chosen
    base_proximity = 1 - abs(diff) / max_dist
    # 对小于众数的距离加额外偏置(可调整幅度)
    if diff < 0:
        proximity_score = min(base_proximity + 0.05, 1.0)
    else:
        proximity_score = base_proximity
    
    # 加权求和:频率权重60%,接近度40%
    score = (frequency_score * 0.6) + (proximity_score * 0.4)
    # 确保评分在0-1范围内
    score = max(min(score, 1.0), 0.0)
    
    print(f"score is {round(score, 3)}...for distance {race_dist}")
    return score

# 测试用例
test_list = [10, 5,5,5,5, 10, 16.09, 10, 10.7, 10,
             10, 5, "Marathon", 15, 10, 10, 5, 10,
             15, 10, 10,10, 5, 5, 5, 5, 5]
calc_racedistance_likeability_score(9, test_list)   # 评分高于11
calc_racedistance_likeability_score(11, test_list)
calc_racedistance_likeability_score(5, test_list)
calc_racedistance_likeability_score(10, test_list)

关键细节说明

  • 非数值处理:将"Marathon"映射为标准马拉松距离,避免统计逻辑出错
  • 偏向调整:通过给小于众数的距离加0.05的额外分,确保9的评分高于11,偏置幅度可根据需求修改
  • 权重分配:频率占更高权重,保证出现次数多的距离即使接近度稍差,评分依然更优
  • 归一化:用列表最大距离作为接近度计算的分母,避免因距离范围过大导致评分失真,同时限制最终评分在0-1区间

内容的提问来源于stack exchange,提问作者dexta

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 14:12:57