DataFrame分组后Timedelta列求均值报错求助
Let's work through your problem step by step, addressing both the empty output and the aggregation error:
1. Why is your output length 0?
This almost always means your intermediate time_to_rent DataFrame has no rows. Let's trace back to the source:
- First, validate if your initial filter returns any data:
If this printsfiltered_payments = user_payments[user_payments.rentComplete] print(filtered_payments.shape) # Should show (number_of_rows, column_count)(0, X), your filter isn't capturing any data. Double-check therentCompletecolumn:- Is it a boolean column? If it uses string values (like 'Yes'/'No' or 'True'/'False'), adjust your filter to match, e.g.,
user_payments[user_payments.rentComplete == 'True']. - It’s also possible there are no rows where
rentCompleteis truthy at all—you’ll need to verify your raw dataset here.
- Is it a boolean column? If it uses string values (like 'Yes'/'No' or 'True'/'False'), adjust your filter to match, e.g.,
2. Fixing the DataError: No numeric types to aggregate
The error occurs because rent_time is a timedelta64 type, and np.mean doesn’t handle this type as seamlessly as pandas’ native aggregation tools. Here are two reliable fixes:
Option 1: Use pandas' built-in mean instead of np.mean
Replace your aggregation line with this:
average_per_user = time_to_rent.groupby('creditCardId').agg({'rent_time': 'mean'})
Pandas natively supports averaging timedelta values, so this will return a timedelta result (e.g., 0 days 02:30:00 for an average of 2.5 hours).
Option 2: Convert timedelta to numeric seconds first
If you prefer working with numeric values, convert rent_time to total seconds before aggregating:
# Convert timedelta to a numeric column (total seconds) time_to_rent['rent_time_seconds'] = time_to_rent['rent_time'].dt.total_seconds() # Now np.mean works perfectly with the numeric type average_per_user = time_to_rent.groupby('creditCardId').agg({'rent_time_seconds': np.mean})
You can convert the result back to a timedelta later if needed:
average_per_user['average_rent_time'] = pd.to_timedelta(average_per_user['rent_time_seconds'], unit='s')
Final Checklist
- Confirm your filtered
user_paymentshas rows before proceeding to groupby operations. - Use one of the aggregation fixes above to avoid the DataError.
内容的提问来源于stack exchange,提问作者OcMaRUS

