如何按type分组为DataFrame中idx>0的行填充上一行val作为prev_val?
Optimal Solution Using Pandas Groupby and Shift
The most efficient and clean way to handle this task is by using pandas' built-in groupby() and shift() functions. These vectorized operations are optimized for speed (even with large datasets) and avoid the slow, messy row-wise loops you might be tempted to write.
How It Works:
- Group by the
typecolumn: This ensures we process each category independently, so we don't accidentally pull values from a different group. - Shift the
valcolumn within each group: Usingshift(1)moves every value in thevalcolumn down by one position inside its group. The first entry in each group becomesNaN, which is exactly what we need for rows whereidx=0(since there's no prior value). - Replace the
prev_valcolumn: Assign the shifted values directly back to theprev_valcolumn to overwrite the originalNaNs where appropriate.
Code Implementation:
import pandas as pd # Your input DataFrame df = pd.DataFrame({ 'idx': [0,1,2,0,1,0,1,2,3], 'prev_val': [pd.NA]*9, 'val': [8,9,7,3,1,2,7,5,4], 'type': ['a','a','a','b','b','c','c','c','c'] }) # The magic line to update prev_val df['prev_val'] = df.groupby('type')['val'].shift(1) print(df)
Output:
idx prev_val val type 0 0 NaN 8 a 1 1 8 9 a 2 2 9 7 a 3 0 NaN 3 b 4 1 3 1 b 5 0 NaN 2 c 6 1 2 7 c 7 2 7 5 c 8 3 5 4 c
Why this is the best approach:
- Speed: Vectorized operations run in optimized C code under the hood, making them way faster than looping through rows manually.
- Simplicity: One line of code does the entire job—easy to read and maintain.
- Reliability: Uses standard pandas functions that are well-tested and widely used in the data community.
内容的提问来源于stack exchange,提问作者Cranjis
相关产品推荐
相关产品推荐

