You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python pandas实现两个DataFrame按邻近值匹配完成列相减

实现代码与逻辑说明

核心逻辑按需求优先级实现,先做快速查询映射避免重复遍历DataFrame,再逐层向外按偏移优先级查找匹配值:

  1. 把bias的第0列、第1列转为键值对字典,查询时间复杂度从O(n)降到O(1)
  2. 遍历missingDateUnique里的每个目标值,偏移量k从1开始递增:每个k优先查i+k是否存在,存在就取对应值;不存在再查i-k,存在就取值;都不存在则k加1继续查找
  3. 把找到的匹配值映射回missingDate的每一行,减去该行第1列原始值得到差值
import pandas as pd

# 示例数据
missingDateUnique = pd.Series({0: 2459650, 9: 2459654})
missingDate = pd.DataFrame(
    {0: [2459650, 2459650,2459650,2459654,2459654,2459654], 
     1: [10, 10,10,14,14,14]},
    index=[0,1,2,9,10,11]
)
bias = pd.DataFrame(
    {0: [2459651, 2459652,2459653,2459655,2459656,2459658,2459659], 
     1: [11, 12,13,15,16,18,19]}
)

# 构建bias快速查询映射
bias_map = dict(zip(bias[0], bias[1]))
match_res = {}

for target in missingDateUnique:
    offset = 1
    while True:
        # 优先匹配正偏移
        pos_candidate = target + offset
        if pos_candidate in bias_map:
            match_res[target] = bias_map[pos_candidate]
            break
        # 正偏移无结果再匹配负偏移
        neg_candidate = target - offset
        if neg_candidate in bias_map:
            match_res[target] = bias_map[neg_candidate]
            break
        # 都无结果则扩大偏移量
        offset += 1

# 计算最终差值
result = missingDate[0].map(match_res) - missingDate[1]
print(result)

运行输出完全符合预期,所有行差值均为1:

0     1
1     1
2     1
9     1
10    1
11    1
dtype: int64

内容的提问来源于stack exchange,提问作者Pritam Das

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 07:45:35