You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用linspace生成列执行Pandas Merge为何出现意外NaN值?

Pandas左连接出现大量NaN的原因及解决办法

问题根源:浮点数精度误差

出现大量NaN的核心原因是浮点数的精确匹配失败:

  • 你手动定义的dfla['score']是直接输入的十进制浮点数(比如-1.9),而np.linspace(-2, 2, 41)生成的数值是通过算法计算得到的浮点数。
  • 计算机存储浮点数时,部分十进制小数无法被二进制精确表示,会产生细微的精度偏差(比如linspace生成的-1.9实际可能是-1.8999999999999999,而你手动输入的是-1.9000000000000001)。
  • Pandas的merge是严格的精确值匹配,只要两个数值的二进制表示不完全一致,就无法匹配,最终左连接后ext_rank列就会出现NaN。

你可以用这段代码验证差异:

# 检查第一个不匹配的数值
print(df_score.loc[1, 'score'] == dfla.loc[1, 'score'])  # 输出False
print(f"df_score的-1.9: {df_score.loc[1, 'score']}")
print(f"dfla的-1.9: {dfla.loc[1, 'score']}")

解决办法

针对这个问题,有几种简单有效的修复方式:

1. 对score列四舍五入

将两个DataFrame的score列统一四舍五入到1位小数,消除精度差异:

dfla['score'] = dfla['score'].round(1)
df_score['score'] = df_score['score'].round(1)
result = pd.merge(df_score, dfla, on="score", how='left')

2. 使用merge_asof(适合有序数据)

由于你的score列是严格递增的有序数据,可以用merge_asof进行近似匹配,指定极小的容忍度:

result = pd.merge_asof(df_score, dfla, on="score", tolerance=1e-9)

3. 转换为Decimal类型(高精度匹配)

如果需要更严谨的精度控制,可以将数值转换为Decimal类型:

from decimal import Decimal

dfla['score'] = dfla['score'].apply(lambda x: Decimal(str(x)))
df_score['score'] = df_score['score'].apply(lambda x: Decimal(str(x)))
result = pd.merge(df_score, dfla, on="score", how='left')

内容的提问来源于stack exchange,提问作者Chris SH

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 11:42:03