在Pandas中实现多ID与单ID场景下的滑动窗口时长计算
Solution for Handling Both Single-ID and Multi-ID Duration Calculation
Got it, let's adjust your code to seamlessly handle both single-ID and multi-ID scenarios in your DataFrame. The core issue with your original code is that when all rows share the same ID, the diff_ids filter only keeps the first row, and the shifted end value ends up as NaT—which doesn't give us the total duration we need.
Step-by-Step Fix
Here's the modified code that checks for unique IDs first, then applies the right logic for each case:
import pandas as pd df = pd.read_csv('df.csv') df['date'] = pd.to_datetime(df['date']) # Check if the DataFrame contains only one unique ID if df['id'].nunique() == 1: # Single ID scenario: calculate total duration from first to last timestamp # Grab the first row to keep consistent columns result_row = df.iloc[0].copy() # Set start to the first timestamp, end to the last timestamp in the DataFrame result_row['start'] = result_row['date'] result_row['end'] = df.iloc[-1]['date'] result_row['duration'] = result_row['end'] - result_row['date'] # Convert the single row to a DataFrame result = pd.DataFrame([result_row]) else: # Multi ID scenario: reuse your original logic (with an optional tweak for the last row) diff_ids = df['id'] != df['id'].shift(1) result = df[diff_ids].copy() result['start'] = result['date'] # Optional: Fill the last row's end with the final timestamp from the original DataFrame # (instead of leaving it as NaT) result['end'] = result['date'].shift(-1).fillna(df.iloc[-1]['date']) result['duration'] = result['end'] - result['start'] print(result)
How It Works
- Single-ID Case: We directly take the first row's data, set
startto the earliest timestamp,endto the latest timestamp in the entire DataFrame, then compute the total duration between them. This gives you the exact output you requested for single-ID scenarios. - Multi-ID Case: We keep your original logic to detect ID changes, but added an optional
fillnato replace the final row'sNaTwith the last timestamp from the original DataFrame—so even the final ID segment gets a valid duration calculation. If you prefer to keep theNaTfor the last row, just remove the.fillna(df.iloc[-1]['date'])part.
Test Results
For your single-ID sample DataFrame, the output will be:
id date value start end duration 0 2 2012-01-01 00:09:45 1 2012-01-01 00:09:45 2012-01-01 00:30:45 00:21:10
For your multi-ID sample, the output (with the optional fill) will have a valid duration for the final row:
id date value start end duration 0 1 2012-01-01 00:09:45 1 2012-01-01 00:09:45 2012-01-01 00:09:55 00:00:10 1 2 2012-01-01 00:09:55 1 2012-01-01 00:09:55 2012-01-01 00:30:20 00:20:25 2 3 2012-01-01 00:30:20 1 2012-01-01 00:30:20 2012-01-01 00:30:45 00:00:25 3 1 2012-01-01 00:30:45 1 2012-01-01 00:30:45 2012-01-01 00:30:45 00:00:00
内容的提问来源于stack exchange,提问作者GKC
相关产品推荐
相关产品推荐

