You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python DataFrame中如何按规则配对application与calculation并合并时间

解决DataFrame中按分组配对时间记录的问题

问题根源

你之前的错误在于仅按Status分组生成配对序号,没有将Transaction group和Application number纳入分组维度。这会导致不同业务组的记录被统一编号,最终在pivot时出现配对错位和大量NaT。

正确解决方案

步骤1:确保时间字段格式正确

先把Created time转换为datetime类型(如果还未转换):

import pandas as pd

df['Created time'] = pd.to_datetime(df['Created time'])

步骤2:生成分组内的配对序号

按Transaction group、Application number、Status三个维度分组,对每组内的记录按时间排序后生成序号,这样同业务组同状态的记录会按时间顺序得到唯一的配对编号:

# 先按业务组、状态、时间排序,保证序号按时间顺序生成
df_sorted = df.sort_values(by=['Transaction group', 'Application number', 'Status', 'Created time'])

# 生成组内序号(从1开始)
df_sorted['pairs'] = df_sorted.groupby(['Transaction group', 'Application number', 'Status']).cumcount() + 1

步骤3:转置或合并生成配对结果

有两种方式可以得到最终结果:

方式一:使用pivot转置

# 按业务组、应用编号、配对序号作为索引,转置Status列
result = df_sorted.pivot(
    index=['Transaction group', 'Application number', 'pairs'],
    columns='Status',
    values='Created time'
).reset_index()

# 重命名列名,符合需求
result = result.rename(columns={'application': 'start_time', 'calculation': 'stop_time'})

# 调整列顺序(可选)
result = result[['pairs', 'Transaction group', 'Application number', 'start_time', 'stop_time']]

方式二:拆分后合并(更灵活,适合状态记录数不一致的场景)

如果同一业务组内application和calculation的记录数量不一致,用merge的方式可以只保留两边都存在的配对:

# 拆分出两种状态的DataFrame
app_df = df_sorted[df_sorted['Status'] == 'application'].rename(columns={'Created time': 'start_time'})
calc_df = df_sorted[df_sorted['Status'] == 'calculation'].rename(columns={'Created time': 'stop_time'})

# 按业务组、应用编号、配对序号合并
result = pd.merge(
    app_df[['Transaction group', 'Application number', 'pairs', 'start_time']],
    calc_df[['Transaction group', 'Application number', 'pairs', 'stop_time']],
    on=['Transaction group', 'Application number', 'pairs'],
    how='inner'
)

# 调整列顺序
result = result[['pairs', 'Transaction group', 'Application number', 'start_time', 'stop_time']]

结果说明

最终的resultDataFrame会包含:

  • pairs:同业务组内的配对编号(第n早的application对应第n早的calculation)
  • Transaction group、Application number:原分组字段
  • start_time:对应application记录的时间
  • stop_time:对应calculation记录的时间

内容的提问来源于stack exchange,提问作者zukowski2012

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 08:20:35