You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Pandas中实现多ID与单ID场景下的滑动窗口时长计算

Solution for Handling Both Single-ID and Multi-ID Duration Calculation

Got it, let's adjust your code to seamlessly handle both single-ID and multi-ID scenarios in your DataFrame. The core issue with your original code is that when all rows share the same ID, the diff_ids filter only keeps the first row, and the shifted end value ends up as NaT—which doesn't give us the total duration we need.

Step-by-Step Fix

Here's the modified code that checks for unique IDs first, then applies the right logic for each case:

import pandas as pd

df = pd.read_csv('df.csv')
df['date'] = pd.to_datetime(df['date'])

# Check if the DataFrame contains only one unique ID
if df['id'].nunique() == 1:
    # Single ID scenario: calculate total duration from first to last timestamp
    # Grab the first row to keep consistent columns
    result_row = df.iloc[0].copy()
    # Set start to the first timestamp, end to the last timestamp in the DataFrame
    result_row['start'] = result_row['date']
    result_row['end'] = df.iloc[-1]['date']
    result_row['duration'] = result_row['end'] - result_row['date']
    # Convert the single row to a DataFrame
    result = pd.DataFrame([result_row])
else:
    # Multi ID scenario: reuse your original logic (with an optional tweak for the last row)
    diff_ids = df['id'] != df['id'].shift(1)
    result = df[diff_ids].copy()
    result['start'] = result['date']
    # Optional: Fill the last row's end with the final timestamp from the original DataFrame
    # (instead of leaving it as NaT)
    result['end'] = result['date'].shift(-1).fillna(df.iloc[-1]['date'])
    result['duration'] = result['end'] - result['start']

print(result)

How It Works

  • Single-ID Case: We directly take the first row's data, set start to the earliest timestamp, end to the latest timestamp in the entire DataFrame, then compute the total duration between them. This gives you the exact output you requested for single-ID scenarios.
  • Multi-ID Case: We keep your original logic to detect ID changes, but added an optional fillna to replace the final row's NaT with the last timestamp from the original DataFrame—so even the final ID segment gets a valid duration calculation. If you prefer to keep the NaT for the last row, just remove the .fillna(df.iloc[-1]['date']) part.

Test Results

For your single-ID sample DataFrame, the output will be:

id                date  value               start                 end duration
0   2 2012-01-01 00:09:45      1 2012-01-01 00:09:45 2012-01-01 00:30:45 00:21:10

For your multi-ID sample, the output (with the optional fill) will have a valid duration for the final row:

id                date  value               start                 end duration
0   1 2012-01-01 00:09:45      1 2012-01-01 00:09:45 2012-01-01 00:09:55 00:00:10
1   2 2012-01-01 00:09:55      1 2012-01-01 00:09:55 2012-01-01 00:30:20 00:20:25
2   3 2012-01-01 00:30:20      1 2012-01-01 00:30:20 2012-01-01 00:30:45 00:00:25
3   1 2012-01-01 00:30:45      1 2012-01-01 00:30:45 2012-01-01 00:30:45 00:00:00

内容的提问来源于stack exchange,提问作者GKC

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 22:47:36