Python分组数据时按规则保留status列值的实现问询
Let's adjust your existing grouping code to include the status logic you need. Here's how to do it step by step:
1. Update the GroupBy Aggregation
Your original code calculates the min createdAt and max updatedAt per group. We'll add an aggregation for the status column (based on your rule: if any row in the group has rentComplete=False, the group status is False; otherwise True).
We can use all() here because it returns True only if all values in the group are True—exactly matching your requirement.
import numpy as np # Group by rentId and creditCardId, with all required aggregations time_to_rent = user_payments.groupby(['rentId', 'creditCardId']).agg( createdAt=('createdAt', np.min), updatedAt=('updatedAt', np.max), status=('rentComplete', lambda x: all(x)) # Core logic for status )
2. Calculate Rental Time Interval
Keep your existing code to compute the time difference:
time_to_rent['rent_time'] = time_to_rent['updatedAt'] - time_to_rent['createdAt']
3. Reorder Columns to Match Your Expected Output
To get the exact column order you showed, reset the index (to bring rentId and creditCardId back as columns) and reorder:
# Reset index and rearrange columns time_to_rent = time_to_rent.reset_index()[['createdAt', 'updatedAt', 'rent_time', 'rentId', 'creditCardId', 'status']]
Final Result
Running this will give you a DataFrame exactly like your expected output:
| createdAt | updatedAt | rent_time | rentId | creditCardId | status |
|---|---|---|---|---|---|
| 2020-09-27 08:44:13.431 | 2020-09-27 09:13:45.675 | 0 days 00:29:32.244333 | 637.0 | 505 | False |
| 2020-09-27 09:14:27.188 | 2020-09-27 12:51:03.394 | 0 days 03:36:36.205525 | 635.0 | 505 | True |
Quick Note
I used rentComplete in the aggregation because that's the boolean column in your sample data (your question mentions status as the boolean column, but your data shows status as 'succeeded'—just adjust the column name here if you meant a different field!).
内容的提问来源于stack exchange,提问作者OcMaRUS

