You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas:相邻C1区间内C_rows重复值的向量化查找

Efficient Vectorized Approach to Find Duplicate C_rows Between Consecutive C1 Entries

Got it, let's solve this efficiently without any slow for loops—perfect for your 50k-row dataset where each interval between C1s has ~15 rows. Here's the step-by-step breakdown:

Step 1: Group Rows by Intervals Between C1s

First, we need to assign a unique group ID to every set of rows between two consecutive C1 entries. We can do this with a cumulative sum on a boolean mask that flags C1 rows:

import pandas as pd

# Your original DataFrame
df = pd.DataFrame()
df['C_rows'] = ['C1', 'C2', 'C3', 'C2', 'C1', 'C2', 'C3', 'C1', 'C2', 'C3', 'C4', 'C1']
df['values'] = ['customer1', 4321, 1266, 5671, 'customer2', 123, 7344,'customer3', 4321, 4444, 5674, 'customer4']

# Create group IDs for each interval between C1s
df['group_id'] = df['C_rows'].eq('C1').cumsum()

What this does: Every time we hit a C1, the cumulative sum increments, so all rows between that C1 and the next one get the same group_id. For example, rows 0-3 (from customer1 to the row before customer2) get group_id=1, rows 4-6 get group_id=2, and so on.

Step 2: Flag Duplicate C_rows Within Each Group

Next, we use groupby combined with transform to check for duplicates in C_rows within each group. We'll use duplicated(keep=False) to mark all instances of a duplicate (not just the second and subsequent ones):

# Flag duplicates within each group
df['is_duplicate_in_group'] = df.groupby('group_id')['C_rows'].transform(lambda x: x.duplicated(keep=False))

# Optional: Exclude C1 rows from duplicate checks (since each group only has one C1)
df['is_duplicate_in_group'] = df['is_duplicate_in_group'] & ~df['C_rows'].eq('C1')

Let's Check the Result

If you print the DataFrame now, you'll see that rows 1 and 3 (both C2 in the first group) are marked as True for is_duplicate_in_group, which is exactly what we want. All other non-duplicate C_rows entries stay False.

Why This Works Efficiently

  • No for loops: All operations use pandas' optimized vectorized functions, which are way faster than looping through rows, especially for large datasets.
  • Groupby overhead is minimal: Since each group only has ~15 rows, the groupby operation is lightning fast even for 50k rows.

内容的提问来源于stack exchange,提问作者A H

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 03:48:51