DataFrame重采样(10ms平均)与缺失行插值处理技术求助
Got it, let's break down exactly how to implement your requirements using pandas. You're spot-on that resample and interpolate are the core tools here—here's a step-by-step implementation that aligns with your example:
Step 1: Prepare the DataFrame
First, we need to convert your timestamp column to a datetime type and set it as the DataFrame index (required for time-based resampling):
import pandas as pd # Sample input data (matches your example) data = { 'ephoc_as_datatime': [ '2000-01-01 11:22:37.130', '2000-01-01 11:22:37.138', '2000-01-01 11:22:37.149', '2000-01-01 11:22:37.156', '2000-01-01 11:22:37.165', '2000-01-01 11:22:37.168', '2000-01-01 11:22:37.169', '2000-01-01 11:22:37.567', '2000-01-01 11:22:38.120' ], 'att1': [0, 1, 2, 3, 4, 5, 3, 7, 8], 'att2': [4, 5, 6, 7, 8, 9, 7, 3, 4] } df = pd.DataFrame(data) # Convert timestamp column to datetime and set as index df['ephoc_as_datatime'] = pd.to_datetime(df['ephoc_as_datatime']) df = df.set_index('ephoc_as_datatime')
Step 2: Resample to 10ms Intervals with Averaging
Next, we resample the data to 10-millisecond intervals, calculating the average value for each interval. The label='right' parameter ensures the index uses the end of each 10ms window (which matches your example output, where the 138ms sample maps to the 140ms timestamp):
# Resample to 10ms intervals, compute mean for each window resampled_df = df.resample('10ms', label='right').mean()
'10ms': Specifies the 10-millisecond resampling intervallabel='right': Sets the index to the end time of each resampling window (omit this if you prefer the start time)mean(): Aggregates all data points within each window into their average
Step 3: Interpolate Missing Values
Finally, we fill in any missing rows (where there were no original data points in a 10ms window) using time-aware interpolation. This method uses the actual time gaps between points to compute linear interpolations, which is more accurate for time series than generic linear interpolation:
# Fill missing values with time-based linear interpolation filled_df = resampled_df.interpolate(method='time') # Optional: Reset index to get the timestamp back as a column filled_df = filled_df.reset_index().rename(columns={'index': 'ephoc_as_datatime'})
method='time': Interpolates based on the time difference between adjacent valid points (ideal for time series data)- If you need a different interpolation method (e.g., quadratic spline), replace
'time'with'quadratic'or another valid pandas interpolation method
Output Verification
When you run this code, you'll get a DataFrame that matches your example:
- The 11:22:37.170 timestamp will have
att1=4andatt2=8(average of the 165ms, 168ms, and 169ms samples) - All gaps between valid resampled points will be filled with interpolated values
- The final timestamp (11:22:38.120) will retain its original value since it falls exactly on a 10ms interval
内容的提问来源于stack exchange,提问作者Carlo Allocca

