如何在Python DataFrame中筛选符合指定事件顺序的行?
event1→event2→event3的行 Hey there! I see the issue with your original code—let's break this down and fix it step by step.
Why your initial approach didn't work
The code df[df['Event'].isin(['event1', 'event2', 'event3'])] only checks if each row's Event value is in the provided list. Since every row in your sample data has one of those three events, it returns the entire DataFrame. This method doesn't care about the order or sequence of events, which is what you actually need.
Solution: Identify complete event1→event2→event3 sequences
First, let's clean up your sample DataFrame by adding column names (this makes everything easier to read):
import pandas as pd df = pd.DataFrame( [ ['event1','01:22:52.134'], ['event2','03:21:31.123'], ['event1','21:12:52.544'], ['event3','23:12:31.216'], ['event1','10:22:02.134'], ['event2','12:21:31.456'], ['event3','14:12:31.789'] ], columns=['Event', 'Timestamp'] )
Now follow these steps to get your desired rows:
- Sort the data by timestamp
Events are ordered by time, so we need to make sure the DataFrame is sorted correctly. First convert theTimestampcolumn to a datetime type, then sort:
df['Timestamp'] = pd.to_datetime(df['Timestamp'], format='%H:%M:%S.%f') df = df.sort_values('Timestamp').reset_index(drop=True)
- Create group identifiers for potential sequences
We'll assign a unique group number every time we encounterevent1—this helps us group together events that start withevent1and follow it:
df['group'] = df['Event'].eq('event1').cumsum()
- Filter groups that match the exact sequence
Usegroupbyto check each group's event order, and keep only groups that exactly match['event1', 'event2', 'event3']:
def validate_sequence(group): return list(group['Event']) == ['event1', 'event2', 'event3'] valid_sequences = df.groupby('group').filter(validate_sequence)
Now valid_sequences will contain only the rows that form a complete, ordered event1→event2→event3 sequence. For your sample data, this will be the three rows with timestamps 10:22:02.134, 12:21:31.456, and 14:12:31.789.
Bonus: Handle edge cases
If your data has duplicate events or incomplete sequences (like event1→event1→event2), this method will still correctly filter out only the full, ordered triplets.
内容的提问来源于stack exchange,提问作者Snowfire777

