如何在Pandas中按条件找出最大滚动窗口并返回对应记录
Got it, let's fix this! You've already figured out how to calculate the length of the longest continuous activity segment (where adjacent intervals are ≤ your specified seconds) for each person. Now we just need to extend that logic to grab the actual records—including their id values—for those longest segments.
Here's the Step-by-Step Solution:
First, instead of only processing the activitydate column in your groupby, we'll work with the full subset of data for each person. This way we can track which records belong to the longest segment.
1. Define a Function to Extract the Longest Segment
This function takes a person's grouped data and your max allowed interval, then returns the full records of their longest valid continuous segment:
def get_longest_segment(group, max_seconds): # Sort the group by activity date to ensure chronological order sorted_group = group.sort_values('activitydate').reset_index(drop=True) # Calculate time difference (in seconds) between consecutive activities sorted_group['time_diff'] = sorted_group['activitydate'].diff().dt.total_seconds().fillna(0) # Assign a unique ID to each continuous segment: # Increment the segment ID whenever the time difference exceeds max_seconds sorted_group['segment_id'] = (sorted_group['time_diff'] > max_seconds).cumsum() # Count the number of records in each segment segment_sizes = sorted_group.groupby('segment_id').size() if segment_sizes.empty: return pd.DataFrame() # Edge case: no records for the person # Find the ID of the longest segment (pick first if multiple segments have same max length) longest_segment_id = segment_sizes.idxmax() # Return all records from the longest segment return sorted_group[sorted_group['segment_id'] == longest_segment_id]
2. Apply the Function to Your Data
Now we'll group your DataFrame by person and apply the function to get the full records for each person's longest segment:
import pandas as pd import datetime import random # Your existing data initialization code (unchanged) def rand_date(x, y): return datetime.datetime(2018, 9, 20, 5, x, y) random.seed(123) dates = [rand_date(random.randint(0, 59), random.randint(0, 59)) for x in range(100)] person = [random.choice(['005', '088']) for x in range(100)] ids = range(100) df = pd.DataFrame({'id':ids, 'person':person, 'activitydate':dates}) # Set your max allowed interval in seconds max_seconds = 50 # Process each person's group and collect results longest_segments = [] for person_id, group in df.groupby('person'): segment = get_longest_segment(group, max_seconds) longest_segments.append(segment) # Combine all segments into a single DataFrame final_result = pd.concat(longest_segments).reset_index(drop=True) # Check the result print(final_result[['id', 'person', 'activitydate']])
How This Works:
- Sorting: We first sort each person's records by
activitydateto ensure we're checking chronological intervals. - Time Difference Calculation:
diff().dt.total_seconds()gives us the number of seconds between consecutive activities; we fill the first value with 0 since there's no prior activity. - Segment ID Assignment: Using
cumsum()on the conditiontime_diff > max_secondscreates unique IDs for each continuous segment. Every time the interval exceeds your limit, a new segment starts. - Extract Longest Segment: We count records per segment, find the longest one, and return all its records (including
id).
Verify Against Your Original Lengths:
If you want to confirm the segment lengths match your original dfg results:
# Check lengths per person print(final_result.groupby('person').size())
You'll see person 005 has 4 records and 088 has 8, which matches your initial calculation.
内容的提问来源于stack exchange,提问作者Matt W.

