如何为Pandas DataFrame连续0/1分组分配连续编号?
Solution for Consecutive Group Numbering in Pandas
Hey there! The key here is to use pandas' vectorized operations to detect when a group changes and assign cumulative IDs to each consecutive group—no slow row-wise loops required.
Step-by-Step Breakdown:
- Spot group transitions: Compare each value in
col1with the previous value usingshift(). A mismatch means we’ve started a new group. - Generate sequential group IDs: Take the cumulative sum of these transition indicators. Each time a change is detected, the sum increments, giving a unique ID for each continuous group.
Code Implementation:
import pandas as pd # Your sample data d = {'col1': [1,1,1,0,0,1,1,0,0,0,1,1]} df1 = pd.DataFrame(data=d) # Add the group number column df1['col2'] = (df1['col1'] != df1['col1'].shift()).cumsum() print(df1)
Expected Output:
col1 col2 0 1 1 1 1 1 2 1 1 3 0 2 4 0 2 5 1 3 6 1 3 7 0 4 8 0 4 9 0 4 10 1 5 11 1 5
Why This Works:
df1['col1'].shift()shifts all values down by one, letting us compare each row to the one before it.- The boolean series
df1['col1'] != df1['col1'].shift()returnsTruewhenever the value changes (including the first row, sinceshift()returnsNaNthere). cumsum()treatsTrueas 1 andFalseas 0, creating a running total that increments exactly when a new group starts.
This method is fully vectorized, so it’s blazingly fast even for large datasets—far more efficient than any loop-based approach.
内容的提问来源于stack exchange,提问作者Gev_2000
相关产品推荐
相关产品推荐

