Pandas按日期与水果分组保留首值其余置零的实现需求
Hey there! Since you're working with a 3M+ row dataset, we need a vectorized, high-performance approach to avoid slow loops. Here's how to solve your problem step by step:
Problem Recap
You have a Pandas DataFrame with a datetime index, and you want to:
- Group by date (index) and Fruit
- Keep the original
Quantityvalue only for the first row in each group - Set
Quantityto 0 for all other rows in the group
Solution Code
This uses vectorized Pandas operations (way faster than row-wise apply for big data):
# Create a boolean mask marking the first row of each (date, Fruit) group first_in_group_mask = df.groupby([df.index, 'Fruit']).cumcount() == 0 # Update Quantity: keep original value where mask is True, set to 0 otherwise df['Quantity'] = df['Quantity'].where(first_in_group_mask, 0)
How It Works
groupby([df.index, 'Fruit']): Groups the DataFrame by both the datetime index (date) and theFruitcolumn.cumcount() == 0: Generates a boolean series whereTrueindicates the first row in each group (sincecumcount()starts counting from 0 for every group).where(): Pandas' vectorized function that retains the originalQuantityvalue when the mask isTrue, and replaces it with 0 when the mask isFalse.
Performance Note
This approach is optimized for large datasets because it avoids row-wise operations (like apply(axis=1) which is slow for millions of rows). Vectorized operations leverage Pandas' underlying C-based optimizations, so it should handle your 3M rows quickly.
Example Output
Running this on your sample data will produce exactly the result you're looking for:
Fruit Quantity01/02/10 Apple 4
01/02/10 Apple 0
01/02/10 Pear 7
01/02/10 Grape 8
01/02/10 Grape 0
02/02/10 Apple 2
02/02/10 Fruit 6
02/02/10 Pear 8
02/02/10 Pear 0
内容的提问来源于stack exchange,提问作者MysterioProgrammer91

