You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提取列表中的相近数值并计算其平均值?

提取列表中相近浮点数的解决方案

问题背景

给定列表:

[0.21,5.0,9.0,0.19,0.2,0.1856,0.9,0.14,0.189]

预期输出:

0.19

该列表中存在一组相近数值:0.21、0.2、0.19、0.1856、0.189,平均值为0.1948。需要实现通用逻辑提取这类相近数值,且无法预先假设输出结果。

原有的整数计数逻辑仅能统计完全重复的元素,不适用于浮点数场景(浮点数几乎不会完全相等),代码如下:

def most_frequent(List):
  counter = 0
  num = List[0]
    
  for i in List:
    curr_frequency = List.count(i)
    if(curr_frequency> counter):
        counter = curr_frequency
        num = i

  return num
    
List = [2, 1, 2, 2, 1, 3]
print(most_frequent(List))

输出:

2

解决思路

针对浮点数的相近值提取,核心思路是聚类分析:

  • 先对列表排序,让相近数值自然聚集
  • 设定容差阈值(可固定或动态计算),判断数值是否属于同一组
  • 统计各组的元素数量,找到规模最大的组
  • 从最大组中提取目标值(如中位数、平均值等)

具体实现

固定容差版本

适合已知数据精度场景,比如小数后两位的数值:

def find_most_cluster_value(numbers, tolerance=0.02):
    # 排序使相近值聚集
    sorted_nums = sorted(numbers)
    max_cluster = []
    current_cluster = [sorted_nums[0]]
    
    for num in sorted_nums[1:]:
        # 判断当前数是否在当前聚类的容差范围内
        if abs(num - current_cluster[-1]) <= tolerance:
            current_cluster.append(num)
        else:
            # 更新最大聚类
            if len(current_cluster) > len(max_cluster):
                max_cluster = current_cluster
            current_cluster = [num]
    # 检查最后一个聚类
    if len(current_cluster) > len(max_cluster):
        max_cluster = current_cluster
    
    # 返回聚类的中位数(示例中刚好为预期输出0.19)
    return max_cluster[len(max_cluster)//2]

# 测试
nums = [0.21,5.0,9.0,0.19,0.2,0.1856,0.9,0.14,0.189]
print(find_most_cluster_value(nums))  # 输出0.19

动态容差版本

通过数据的标准差自动计算容差,适配不同分布的数据:

import statistics

def find_most_cluster_value(numbers):
    # 用标准差的30%作为容差(比例可根据场景调整)
    std_dev = statistics.stdev(numbers)
    tolerance = std_dev * 0.3
    
    sorted_nums = sorted(numbers)
    max_cluster = []
    current_cluster = [sorted_nums[0]]
    
    for num in sorted_nums[1:]:
        if abs(num - current_cluster[-1]) <= tolerance:
            current_cluster.append(num)
        else:
            if len(current_cluster) > len(max_cluster):
                max_cluster = current_cluster
            current_cluster = [num]
    if len(current_cluster) > len(max_cluster):
        max_cluster = current_cluster
    
    # 返回聚类中位数
    return max_cluster[len(max_cluster)//2]

# 测试
nums = [0.21,5.0,9.0,0.19,0.2,0.1856,0.9,0.14,0.189]
print(find_most_cluster_value(nums))  # 输出0.19

说明

  • 若需要返回平均值而非中位数,可将返回部分替换为:return round(sum(max_cluster)/len(max_cluster), 2)
  • 容差比例或固定值可根据实际数据的精度、分布灵活调整

内容的提问来源于stack exchange,提问作者Chaitanya Krishna

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 08:34:57