如何在Pandas DataFrame同组内筛选'A'后紧跟'D'的行对
Solution for Finding Consecutive A-D Pairs Within Number Groups
Hey there! Let's work through this problem to find those consecutive A-D pairs that stay strictly within the same Number group. First, I spotted a tiny typo in your sorting code: you used 'number' (lowercase) but your column is named 'Number' (uppercase) — that would throw an error, so we'll fix that first.
Step-by-Step Approach
The core challenge here is to only compare rows within their own Number group. We can use pandas' groupby combined with shift() to check adjacent rows without crossing group boundaries.
Full Code Solution
import pandas as pd # Your original dataset data = [['John', 6, 'A'], ['Paul', 6, 'D'], ['Juli', 9, 'D'], ['Geeta', 4, 'A'], ['Jay', 6, 'D'], ['Sara', 6, 'A'], ['Mario', 3, 'D'], ['Peter', 6, 'A'], ['Jin', 6, 'D'], ['Carl', 6, 'A']] df = pd.DataFrame(data, columns=['Name', 'Number', 'Label']) # 1. Correctly sort by Number and reset index to avoid gaps interfering with row comparisons df_sorted = df.sort_values('Number').reset_index(drop=True) # 2. For each Number group, get the Label from the previous and next row (only within the group) df_sorted['prev_label'] = df_sorted.groupby('Number')['Label'].shift(1) df_sorted['next_label'] = df_sorted.groupby('Number')['Label'].shift(-1) # 3. Create a filter to capture both sides of the A-D pair: # - Rows where Label is 'A' and the next row in the same group is 'D' # - Rows where Label is 'D' and the previous row in the same group is 'A' pair_mask = ((df_sorted['Label'] == 'A') & (df_sorted['next_label'] == 'D')) | \ ((df_sorted['Label'] == 'D') & (df_sorted['prev_label'] == 'A')) # 4. Apply the filter and clean up temporary columns result = df_sorted[pair_mask].drop(columns=['prev_label', 'next_label']) print(result)
Expected Output
Running this code will produce exactly the result you're targeting:
Name Number Label 2 Peter 4 A 3 Jin 4 D 4 John 6 A 5 Paul 6 D
How It Works
- Grouped Shift: Using
groupby('Number')withshift()ensures we only compare rows within the same Number group. Cross-group rows will show up asNaN, so they can't be mistakenly matched. - Dual Mask: We capture both the leading 'A' and trailing 'D' of each valid pair to get the complete row set you need.
- Index Reset: Resetting the index after sorting eliminates gaps from the original data, making the
shift()function work reliably for adjacent row comparisons.
内容的提问来源于stack exchange,提问作者Mario Fernandez
相关产品推荐
相关产品推荐

