You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中按±1阈值对齐并合并3个DataFrame的X列?

解决方案:Pandas实现X列对齐合并

我会用Python和Pandas帮你完成这个DataFrame的转换,完全匹配你描述的规则。先来看具体的实现步骤:

1. 准备原始数据

首先我们把你提供的原始数据构建成Pandas DataFrame:

import pandas as pd
import numpy as np

# 原始数据
data = {
    'X-1': [411.726266, 437.692665, 448.596937, 481.634531, 486.971163, 492.505388, 500.030500, 508.106596, 513.424161, 535.131828, 551.779516, 559.050425, 562.108257, 591.436797],
    'X-1_y': [1387.29, 677.39, 2276.35, 0.00, 0.00, 0.00, 810.45, 0.00, 0.00, 0.00, 3124.92, 1081.89, 3532.11, 0.00],
    'X-2': [437.404307, 448.557534, 481.550490, 486.966310, 492.578155, 500.000000, 508.218825, 513.579177, 515.308245, 534.346333, 551.712654, 559.084859, 562.253910, 590.659433],
    'X-2_y': [3755.08, 1460.70, 0.00, 0.00, 0.00, 608.22, 0.00, 0.00, 0.00, 0.00, 2226.94, 984.05, 3686.94, 0.00],
    'X-3': [437.273585, 448.760155, 481.566018, 487.208899, 492.446192, 500.153040, 508.315935, 513.620953, 515.175867, 534.985459, 551.680943, 559.087271, 562.234223, 591.396752],
    'X-3_y': [3360.85, 981.45, 0.00, 0.00, 0.00, 0.00, 0.00, 9582.45, 0.00, 0.00, 2522.73, 1600.48, 4495.73, 0.00]
}

df = pd.DataFrame(data)

2. 转换为长格式方便分组

我们把原来的宽格式数据转换成长格式,这样每个X值和对应的y值单独成一行,同时记录它来自哪个X列(X-1/X-2/X-3)以及原始行号:

# 转换为长格式
long_data = []
for idx, row in df.iterrows():
    long_data.append({'X': row['X-1'], 'y': row['X-1_y'], 'source': 'X-1', 'original_row': idx})
    long_data.append({'X': row['X-2'], 'y': row['X-2_y'], 'source': 'X-2', 'original_row': idx})
    long_data.append({'X': row['X-3'], 'y': row['X-3_y'], 'source': 'X-3', 'original_row': idx})

long_df = pd.DataFrame(long_data).sort_values('X').reset_index(drop=True)

3. 按X值相似度分组

接下来我们把所有X值按「彼此在±1范围内」的规则分组。这里先对X排序,然后逐个判断是否加入当前组:

# 分组:将X值彼此在±1范围内的归为同一组
groups = []
current_group = [long_df.iloc[0]]

for i in range(1, len(long_df)):
    current_x = long_df.iloc[i]['X']
    # 用当前组的平均值判断,确保新X和组内所有值的差距都在允许范围内
    group_avg = np.mean([item['X'] for item in current_group])
    if abs(current_x - group_avg) <= 1:
        current_group.append(long_df.iloc[i])
    else:
        groups.append(current_group)
        current_group = [long_df.iloc[i]]
groups.append(current_group)  # 别忘了最后一组

4. 生成目标DataFrame

对每个分组,计算平均X值,然后提取三个X列对应的y值(没有对应值就填0):

# 处理每个分组,生成结果行
result_rows = []
for group in groups:
    avg_x = np.mean([item['X'] for item in group])
    # 提取三个y值,找不到就填0
    x1_y = next((item['y'] for item in group if item['source'] == 'X-1'), 0.0)
    x2_y = next((item['y'] for item in group if item['source'] == 'X-2'), 0.0)
    x3_y = next((item['y'] for item in group if item['source'] == 'X-3'), 0.0)
    
    result_rows.append({
        'avg X': round(avg_x, 6),  # 保留6位小数和示例一致
        'X-1_y': x1_y,
        'X-2_y': x2_y,
        'X-3_y': x3_y
    })

# 转成DataFrame
result_df = pd.DataFrame(result_rows).reset_index(drop=True)

运行后你会得到和你预期完全一致的结果:

avg X  X-1_y  X-2_y  X-3_y
0   411.726266  1387.29      0.00      0.00
1   437.456852   677.39   3755.08   3360.85
2   448.638209  2276.35   1460.70    981.45
3   481.583680      0.00      0.00      0.00
4   487.048791      0.00      0.00      0.00
5   492.509912      0.00      0.00      0.00
6   500.061180   810.45    608.22      0.00
7   508.213785      0.00      0.00      0.00
8   513.541430      0.00      0.00   9582.45
9   515.242056      0.00      0.00      0.00
10  534.821206      0.00      0.00      0.00
11  551.724371  3124.92   2226.94   2522.73
12  559.074185  1081.89    984.05   1600.48
13  562.198797  3532.11   3686.94   4495.73
14  591.164327      0.00      0.00      0.00

这个逻辑完美匹配你描述的规则:比如原始第一行的X-1单独成组,而X-2、X-3和第二行的X-1因为值接近被合并成一组取平均,完全符合你给出的示例结果。

内容的提问来源于stack exchange,提问作者sepehr khatir

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 07:42:54