You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

DataFrame计算item_price后unique()返回重复3.99的问题排查

问题描述

现有如下结构的DataFrame:

quantity  price
202     1       3.99
287     1       3.99
309     1       3.99
1345    3       11.97
1681    1       14.36
... ... ...
275754  1       3.59
275922  1       3.99
275927  1       3.99
276012  1       3.99
276065  1       3.59

执行代码计算单品价格:

df['item_price'] = (df['price'] / df['quantity'])

调用df.item_price.unique()后得到结果(注意重复的3.99):

[ 3.99,  0.  ,  3.99,  3.59, 14.36]

尝试过strip()、str.replace(' ', '')等字符串处理方案,均无效,3.99始终重复,求原因及解决办法。


原因分析

这不是字符串格式问题,是浮点数精度误差导致的。

浮点数在计算机中以二进制存储,像3.99这类十进制小数无法被精确表示为二进制浮点数,会存在微小的舍入误差。看起来都是3.99的两个值,实际可能是3.9900000000000002和3.9899999999999998这类极其接近的数值,打印时被格式化显示为3.99,但底层存储值不同,所以unique()会将它们判定为不同值。

你用字符串处理方法无效,是因为这些列的类型是数值型(float),不是字符串,字符串处理函数对数值型数据不起作用。


解决办法
  1. 对浮点数进行四舍五入:
    保留两位小数消除精度误差:

    df['item_price'] = (df['price'] / df['quantity']).round(2)
    df.item_price.unique()
    
  2. 使用Decimal类型精确计算:
    借助decimal模块进行高精度运算,避免浮点数误差:

    from decimal import Decimal, getcontext
    
    getcontext().prec = 4  # 设置精度
    df['item_price'] = df.apply(lambda row: Decimal(str(row['price'])) / Decimal(str(row['quantity'])), axis=1)
    df.item_price.unique()
    
  3. 用numpy的isclose进行近似去重:
    如果不想修改原始数据,可用numpy的近似比较筛选唯一值:

    import numpy as np
    
    prices = df.item_price.values
    unique_prices = []
    for p in prices:
        if not any(np.isclose(p, up) for up in unique_prices):
            unique_prices.append(p)
    print(unique_prices)
    

内容的提问来源于stack exchange,提问作者Ingo Eyring

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 17:03:15