按小时重采样DataFrame:Sms、Call、Internet列取均值实现问题
Got it, let's break down why your two code attempts threw errors and get this resampling working properly.
Why Those Errors Happened
Let's unpack each issue:
- First error: When you ran
df1.reset_index().set_index('TIME').resample('1H').mean(), theTIMEcolumn wasn't converted to a datetime type yet. Even though you set it as the index, it was just a regularIndexobject—not aDatetimeIndex—andresample()only works with time-based indexes. - Second error: You converted
TIMEto datetime, but forgot to set it as the DataFrame's index.resample()operates on the index by default, and your index was still the auto-numberedRangeIndex, hence the mismatch error.
Step-by-Step Solution
Here's how to fix this, with two common scenarios depending on how you want to handle the ID column:
Scenario 1: Resample the entire dataset (hourly averages across all rows)
If you just want overall hourly means for the three target columns:
import pandas as pd # 1. Convert TIME column to datetime (this is critical!) df1['TIME'] = pd.to_datetime(df1['TIME']) # Make sure you use df1 here, not 'data' to avoid variable mix-ups # 2. Set TIME as the DataFrame's index df1 = df1.set_index('TIME') # 3. Resample by hour and calculate mean for your target columns resampled_df = df1[['Sms', 'Call', 'Internet']].resample('1H').mean()
Scenario 2: Resample per ID (hourly averages for each individual ID)
If your ID represents distinct entities (like users or devices) and you need hourly stats per ID:
import pandas as pd # 1. Convert TIME to datetime first df1['TIME'] = pd.to_datetime(df1['TIME']) # 2. Group by ID, then resample each group by hour and compute mean resampled_df = df1.groupby('ID').resample('1H', on='TIME')[['Sms', 'Call', 'Internet']].mean().reset_index()
Note: Using on='TIME' lets you keep TIME as a column instead of setting it as the index, which preserves the original structure if that's what you need.
Quick Troubleshooting Tip
If pd.to_datetime() fails, your TIME column might have a non-standard format. Specify the format explicitly to fix this:
# Adjust the format string to match your actual TIME values (e.g., 'YYYY-MM-DD HH:MM:SS') df1['TIME'] = pd.to_datetime(df1['TIME'], format='%Y-%m-%d %H:%M:%S')
内容的提问来源于stack exchange,提问作者Shruti Bothe

