Python Pandas自定义向下取整函数异常:无法正确取整至步长值
修复Pandas自定义向下取整函数的精度问题
问题描述
自定义的向下取整函数在处理步长为0.1的1.400时,错误地取整为1.3,而预期结果应为1.4。原函数及调用代码如下:
import pandas as pd # 定义自定义取整函数 def custom_round(x, step): return x if round(x / step) * step == x else (x // step) * step # 示例数据 data = { "1": [1.300, 1.400, 1.333, 1.364, 1.400], "X": [5.0, 5.0, 5.0, 4.5, 4.5] } outcome = pd.DataFrame(data) step = 0.1 # 应用函数生成新列 outcome["1_cluster"] = outcome["1"].apply(lambda x: custom_round(x, step)) outcome["X_cluster"] = outcome["X"].apply(lambda x: custom_round(x, step)) print(outcome)
当前错误输出
1 X 1_cluster X_cluster 0 1.300 5.0 1.3 5.0 1 1.400 5.0 1.3 5.0 2 1.333 5.0 1.3 5.0 3 1.364 4.5 1.3 4.5 4 1.400 4.5 1.3 4.5
期望输出
1 X 1_cluster X_cluster 0 1.300 5.0 1.3 5.0 1 1.400 5.0 1.4 5.0 2 1.333 5.0 1.3 5.0 3 1.364 4.5 1.3 4.5 4 1.400 4.5 1.4 4.5
问题根源
原函数中使用round(x / step) * step == x判断是否为步长整数倍,但二进制浮点数无法精确表示部分十进制小数(比如1.4),导致1.4 / 0.1的实际计算结果为13.999999999999998,相等判断返回False,进而执行(x // step) * step,而1.4 // 0.1因浮点数误差得到13.0,最终取整为1.3。
修复方案
方案一:用math.isclose替换直接相等判断
通过允许极小的误差范围,避免浮点数精度导致的误判:
import pandas as pd import math def custom_round(x, step): # 容忍1e-9的相对误差,判断是否为步长整数倍 if math.isclose(x, round(x / step) * step, rel_tol=1e-9): return x return (x // step) * step # 调用代码不变 data = { "1": [1.300, 1.400, 1.333, 1.364, 1.400], "X": [5.0, 5.0, 5.0, 4.5, 4.5] } outcome = pd.DataFrame(data) step = 0.1 outcome["1_cluster"] = outcome["1"].apply(lambda x: custom_round(x, step)) outcome["X_cluster"] = outcome["X"].apply(lambda x: custom_round(x, step)) print(outcome)
方案二:使用decimal模块处理精确小数
彻底规避浮点数精度问题,用十进制精确运算:
import pandas as pd from decimal import Decimal, getcontext def custom_round(x, step): # 设置足够精度确保小数运算准确 getcontext().prec = 10 x_dec = Decimal(str(x)) step_dec = Decimal(str(step)) # 向下取整到最近的步长倍数 multiple = (x_dec / step_dec).to_integral_value(rounding='ROUND_FLOOR') return float(multiple * step_dec) # 调用代码不变 data = { "1": [1.300, 1.400, 1.333, 1.364, 1.400], "X": [5.0, 5.0, 5.0, 4.5, 4.5] } outcome = pd.DataFrame(data) step = 0.1 outcome["1_cluster"] = outcome["1"].apply(lambda x: custom_round(x, step)) outcome["X_cluster"] = outcome["X"].apply(lambda x: custom_round(x, step)) print(outcome)
方案三:转换为整数运算(适合步长为10的负次幂场景)
通过放大数值为整数,避免浮点数误差:
import pandas as pd def custom_round(x, step): # 计算步长的小数位数 step_str = str(step) decimal_places = len(step_str.split('.')[1]) if '.' in step_str else 0 factor = 10 ** decimal_places # 转换为整数后向下取整 x_int = int(round(x * factor)) step_int = int(step * factor) rounded_int = (x_int // step_int) * step_int return rounded_int / factor # 调用代码不变 data = { "1": [1.300, 1.400, 1.333, 1.364, 1.400], "X": [5.0, 5.0, 5.0, 4.5, 4.5] } outcome = pd.DataFrame(data) step = 0.1 outcome["1_cluster"] = outcome["1"].apply(lambda x: custom_round(x, step)) outcome["X_cluster"] = outcome["X"].apply(lambda x: custom_round(x, step)) print(outcome)
以上三种方案均能得到预期的取整结果。
内容的提问来源于stack exchange,提问作者arrabattapp man
相关产品推荐
相关产品推荐

