DataFrame计算item_price后unique()返回重复3.99的问题排查
问题描述
现有如下结构的DataFrame:
quantity price 202 1 3.99 287 1 3.99 309 1 3.99 1345 3 11.97 1681 1 14.36 ... ... ... 275754 1 3.59 275922 1 3.99 275927 1 3.99 276012 1 3.99 276065 1 3.59
执行代码计算单品价格:
df['item_price'] = (df['price'] / df['quantity'])
调用df.item_price.unique()后得到结果(注意重复的3.99):
[ 3.99, 0. , 3.99, 3.59, 14.36]
尝试过strip()、str.replace(' ', '')等字符串处理方案,均无效,3.99始终重复,求原因及解决办法。
原因分析
这不是字符串格式问题,是浮点数精度误差导致的。
浮点数在计算机中以二进制存储,像3.99这类十进制小数无法被精确表示为二进制浮点数,会存在微小的舍入误差。看起来都是3.99的两个值,实际可能是3.9900000000000002和3.9899999999999998这类极其接近的数值,打印时被格式化显示为3.99,但底层存储值不同,所以unique()会将它们判定为不同值。
你用字符串处理方法无效,是因为这些列的类型是数值型(float),不是字符串,字符串处理函数对数值型数据不起作用。
解决办法
对浮点数进行四舍五入:
保留两位小数消除精度误差:df['item_price'] = (df['price'] / df['quantity']).round(2) df.item_price.unique()使用Decimal类型精确计算:
借助decimal模块进行高精度运算,避免浮点数误差:from decimal import Decimal, getcontext getcontext().prec = 4 # 设置精度 df['item_price'] = df.apply(lambda row: Decimal(str(row['price'])) / Decimal(str(row['quantity'])), axis=1) df.item_price.unique()用numpy的isclose进行近似去重:
如果不想修改原始数据,可用numpy的近似比较筛选唯一值:import numpy as np prices = df.item_price.values unique_prices = [] for p in prices: if not any(np.isclose(p, up) for up in unique_prices): unique_prices.append(p) print(unique_prices)
内容的提问来源于stack exchange,提问作者Ingo Eyring
相关产品推荐
相关产品推荐

