如何在Pandas中对Timestamp按两两分组执行行间减法运算
To calculate the time difference between every pair of consecutive rows (grouped two at a time) in a Pandas Series of Timestamps, you can use one of these straightforward approaches:
Method 1: Group by Index Division
This method explicitly groups rows into pairs using integer division of the index, then computes the difference between the second and first element in each group.
import pandas as pd # Create your Timestamp series timestamps = pd.Series([ pd.Timestamp('2020-11-26 20:00:00'), pd.Timestamp('2020-11-26 21:00:00'), pd.Timestamp('2020-11-26 22:00:00'), pd.Timestamp('2020-11-26 23:30:00') ]) # Group rows into pairs (index // 2 creates group IDs: 0,0,1,1) pair_diffs = timestamps.groupby(timestamps.index // 2).apply(lambda x: x.iloc[1] - x.iloc[0]) print(pair_diffs)
Output:
0 0 days 01:00:00 1 0 days 01:30:00 dtype: timedelta64[ns]
Notes:
- This works best when you have an even number of rows. If you have an odd number, the last group will only have one element, and
x.iloc[1]will throw an error. To handle this, filter out groups with fewer than 2 elements first:pair_diffs = timestamps.groupby(timestamps.index // 2).filter(lambda x: len(x) == 2).groupby(timestamps.index // 2).apply(lambda x: x.iloc[1] - x.iloc[0])
Method 2: Shift and Slice
This approach uses shift() to align each timestamp with the next one, computes all consecutive differences, then selects only the differences from your paired groups.
import pandas as pd timestamps = pd.Series([ pd.Timestamp('2020-11-26 20:00:00'), pd.Timestamp('2020-11-26 21:00:00'), pd.Timestamp('2020-11-26 22:00:00'), pd.Timestamp('2020-11-26 23:30:00') ]) # Compute differences between each timestamp and the next one all_diffs = timestamps.shift(-1) - timestamps # Select every other difference starting from index 0 (pairs 0&1, 2&3) pair_diffs = all_diffs.iloc[::2].dropna() print(pair_diffs)
Output:
0 0 days 01:00:00 2 0 days 01:30:00 dtype: timedelta64[ns]
Notes:
shift(-1)moves each value down by one position, sotimestamps.shift(-1)[i]is the next timestamp aftertimestamps[i].iloc[::2]slices the series to take every 2nd element starting at index 0, which gives us exactly the differences between each pair of rows.dropna()removes any NaN values that result from the last row (since there's no next timestamp to subtract against).
Both methods will give you the desired timedelta results. Choose the one that fits your code style and edge case handling needs best!
内容的提问来源于stack exchange,提问作者Zebra125

