Pandas:相邻C1区间内C_rows重复值的向量化查找
Got it, let's solve this efficiently without any slow for loops—perfect for your 50k-row dataset where each interval between C1s has ~15 rows. Here's the step-by-step breakdown:
Step 1: Group Rows by Intervals Between C1s
First, we need to assign a unique group ID to every set of rows between two consecutive C1 entries. We can do this with a cumulative sum on a boolean mask that flags C1 rows:
import pandas as pd # Your original DataFrame df = pd.DataFrame() df['C_rows'] = ['C1', 'C2', 'C3', 'C2', 'C1', 'C2', 'C3', 'C1', 'C2', 'C3', 'C4', 'C1'] df['values'] = ['customer1', 4321, 1266, 5671, 'customer2', 123, 7344,'customer3', 4321, 4444, 5674, 'customer4'] # Create group IDs for each interval between C1s df['group_id'] = df['C_rows'].eq('C1').cumsum()
What this does: Every time we hit a C1, the cumulative sum increments, so all rows between that C1 and the next one get the same group_id. For example, rows 0-3 (from customer1 to the row before customer2) get group_id=1, rows 4-6 get group_id=2, and so on.
Step 2: Flag Duplicate C_rows Within Each Group
Next, we use groupby combined with transform to check for duplicates in C_rows within each group. We'll use duplicated(keep=False) to mark all instances of a duplicate (not just the second and subsequent ones):
# Flag duplicates within each group df['is_duplicate_in_group'] = df.groupby('group_id')['C_rows'].transform(lambda x: x.duplicated(keep=False)) # Optional: Exclude C1 rows from duplicate checks (since each group only has one C1) df['is_duplicate_in_group'] = df['is_duplicate_in_group'] & ~df['C_rows'].eq('C1')
Let's Check the Result
If you print the DataFrame now, you'll see that rows 1 and 3 (both C2 in the first group) are marked as True for is_duplicate_in_group, which is exactly what we want. All other non-duplicate C_rows entries stay False.
Why This Works Efficiently
- No for loops: All operations use pandas' optimized vectorized functions, which are way faster than looping through rows, especially for large datasets.
- Groupby overhead is minimal: Since each group only has ~15 rows, the groupby operation is lightning fast even for 50k rows.
内容的提问来源于stack exchange,提问作者A H

