如何提取列表中的相近数值并计算其平均值?
提取列表中相近浮点数的解决方案
问题背景
给定列表:
[0.21,5.0,9.0,0.19,0.2,0.1856,0.9,0.14,0.189]
预期输出:
0.19
该列表中存在一组相近数值:0.21、0.2、0.19、0.1856、0.189,平均值为0.1948。需要实现通用逻辑提取这类相近数值,且无法预先假设输出结果。
原有的整数计数逻辑仅能统计完全重复的元素,不适用于浮点数场景(浮点数几乎不会完全相等),代码如下:
def most_frequent(List): counter = 0 num = List[0] for i in List: curr_frequency = List.count(i) if(curr_frequency> counter): counter = curr_frequency num = i return num List = [2, 1, 2, 2, 1, 3] print(most_frequent(List))
输出:
2
解决思路
针对浮点数的相近值提取,核心思路是聚类分析:
- 先对列表排序,让相近数值自然聚集
- 设定容差阈值(可固定或动态计算),判断数值是否属于同一组
- 统计各组的元素数量,找到规模最大的组
- 从最大组中提取目标值(如中位数、平均值等)
具体实现
固定容差版本
适合已知数据精度场景,比如小数后两位的数值:
def find_most_cluster_value(numbers, tolerance=0.02): # 排序使相近值聚集 sorted_nums = sorted(numbers) max_cluster = [] current_cluster = [sorted_nums[0]] for num in sorted_nums[1:]: # 判断当前数是否在当前聚类的容差范围内 if abs(num - current_cluster[-1]) <= tolerance: current_cluster.append(num) else: # 更新最大聚类 if len(current_cluster) > len(max_cluster): max_cluster = current_cluster current_cluster = [num] # 检查最后一个聚类 if len(current_cluster) > len(max_cluster): max_cluster = current_cluster # 返回聚类的中位数(示例中刚好为预期输出0.19) return max_cluster[len(max_cluster)//2] # 测试 nums = [0.21,5.0,9.0,0.19,0.2,0.1856,0.9,0.14,0.189] print(find_most_cluster_value(nums)) # 输出0.19
动态容差版本
通过数据的标准差自动计算容差,适配不同分布的数据:
import statistics def find_most_cluster_value(numbers): # 用标准差的30%作为容差(比例可根据场景调整) std_dev = statistics.stdev(numbers) tolerance = std_dev * 0.3 sorted_nums = sorted(numbers) max_cluster = [] current_cluster = [sorted_nums[0]] for num in sorted_nums[1:]: if abs(num - current_cluster[-1]) <= tolerance: current_cluster.append(num) else: if len(current_cluster) > len(max_cluster): max_cluster = current_cluster current_cluster = [num] if len(current_cluster) > len(max_cluster): max_cluster = current_cluster # 返回聚类中位数 return max_cluster[len(max_cluster)//2] # 测试 nums = [0.21,5.0,9.0,0.19,0.2,0.1856,0.9,0.14,0.189] print(find_most_cluster_value(nums)) # 输出0.19
说明
- 若需要返回平均值而非中位数,可将返回部分替换为:
return round(sum(max_cluster)/len(max_cluster), 2) - 容差比例或固定值可根据实际数据的精度、分布灵活调整
内容的提问来源于stack exchange,提问作者Chaitanya Krishna
相关产品推荐
相关产品推荐

