You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何让Pandas merge_asof结果用NaN填充而非重复值?

问题描述

我用以下代码进行时间序列的匹配操作:

import pandas as pd

left = pd.DataFrame({"left_val": [1, 2, 3, 6, 7]}, index=pd.to_datetime([1, 2, 3, 6, 7], unit='s'))
right = pd.DataFrame({"right_val": ["a", "b", "c"]}, index=pd.to_datetime([1, 5, 10], unit='s'))

# 筛选出在left时间区间内的样本
right_filtered = right[(right.index >= left.index.min()) & (right.index <= left.index.max())]

output = pd.merge_asof(left, right_filtered, left_index=True, right_index=True, direction="nearest")

当前得到的输出是:

left_val right_val
1970-01-01 00:00:01         1         a
1970-01-01 00:00:02         2         a
1970-01-01 00:00:03         3         a
1970-01-01 00:00:06         6         b
1970-01-01 00:00:07         7         b

但我期望得到的稀疏输出是:

left_val right_val
1970-01-01 00:00:01         1         a
1970-01-01 00:00:02         2         NaN
1970-01-01 00:00:03         3         NaN
1970-01-01 00:00:06         6         b
1970-01-01 00:00:07         7         NaN

核心需求是让right中的每个值仅在输出DataFrame中出现一次,其余位置用NaN填充以节省空间。我不想遍历结果把重复值设为NaN,原因是:

  • 速度效率问题
  • 如果right中有连续值,这种方法会丢失原始信息

我没找到merge_asof或其他方法的相关参数实现这个需求,特此求助。

解决方案

可以通过定位每个right索引在left中的最近邻索引,然后仅在这些匹配位置填充right的值,其余位置保留NaN。这种方法无需遍历,效率更高,且能保留原始信息。

代码实现如下:

import pandas as pd

left = pd.DataFrame({"left_val": [1, 2, 3, 6, 7]}, index=pd.to_datetime([1, 2, 3, 6, 7], unit='s'))
right = pd.DataFrame({"right_val": ["a", "b", "c"]}, index=pd.to_datetime([1, 5, 10], unit='s'))

# 筛选right中在left时间区间内的样本
right_filtered = right[(right.index >= left.index.min()) & (right.index <= left.index.max())]

# 找到每个right索引对应的left中的最近邻索引位置
nearest_left_positions = left.index.get_indexer(right_filtered.index, method='nearest')
# 获取对应的left索引标签
matched_indices = left.index[nearest_left_positions]

# 创建仅在匹配位置有值的稀疏Series
right_sparse = pd.Series(right_filtered['right_val'].values, index=matched_indices)

# 和left合并,自动为未匹配位置填充NaN
output = left.join(right_sparse, how='left')

运行后得到的输出完全符合预期:

left_val right_val
1970-01-01 00:00:01         1         a
1970-01-01 00:00:02         2         NaN
1970-01-01 00:00:03         3         NaN
1970-01-01 00:00:06         6         b
1970-01-01 00:00:07         7         NaN

原理说明

  1. left.index.get_indexer(right_filtered.index, method='nearest'):利用pandas内置的索引匹配功能,快速定位每个right索引在left索引中的最近邻位置,返回位置数组。
  2. 通过位置数组获取对应的left索引标签,创建仅包含匹配位置的稀疏Series。
  3. 使用join方法合并left和稀疏Series,自动为未匹配的位置填充NaN,生成符合要求的稀疏结果。

这种方法依赖pandas底层优化的索引操作,效率远高于遍历修改,同时完整保留了right的原始信息。

内容的提问来源于stack exchange,提问作者Manu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 13:28:12