You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

DataFrame重采样(10ms平均)与缺失行插值处理技术求助

Solution for Resampling and Interpolating 100Hz Time Series Data

Got it, let's break down exactly how to implement your requirements using pandas. You're spot-on that resample and interpolate are the core tools here—here's a step-by-step implementation that aligns with your example:

Step 1: Prepare the DataFrame

First, we need to convert your timestamp column to a datetime type and set it as the DataFrame index (required for time-based resampling):

import pandas as pd

# Sample input data (matches your example)
data = {
    'ephoc_as_datatime': [
        '2000-01-01 11:22:37.130',
        '2000-01-01 11:22:37.138',
        '2000-01-01 11:22:37.149',
        '2000-01-01 11:22:37.156',
        '2000-01-01 11:22:37.165',
        '2000-01-01 11:22:37.168',
        '2000-01-01 11:22:37.169',
        '2000-01-01 11:22:37.567',
        '2000-01-01 11:22:38.120'
    ],
    'att1': [0, 1, 2, 3, 4, 5, 3, 7, 8],
    'att2': [4, 5, 6, 7, 8, 9, 7, 3, 4]
}

df = pd.DataFrame(data)

# Convert timestamp column to datetime and set as index
df['ephoc_as_datatime'] = pd.to_datetime(df['ephoc_as_datatime'])
df = df.set_index('ephoc_as_datatime')

Step 2: Resample to 10ms Intervals with Averaging

Next, we resample the data to 10-millisecond intervals, calculating the average value for each interval. The label='right' parameter ensures the index uses the end of each 10ms window (which matches your example output, where the 138ms sample maps to the 140ms timestamp):

# Resample to 10ms intervals, compute mean for each window
resampled_df = df.resample('10ms', label='right').mean()
  • '10ms': Specifies the 10-millisecond resampling interval
  • label='right': Sets the index to the end time of each resampling window (omit this if you prefer the start time)
  • mean(): Aggregates all data points within each window into their average

Step 3: Interpolate Missing Values

Finally, we fill in any missing rows (where there were no original data points in a 10ms window) using time-aware interpolation. This method uses the actual time gaps between points to compute linear interpolations, which is more accurate for time series than generic linear interpolation:

# Fill missing values with time-based linear interpolation
filled_df = resampled_df.interpolate(method='time')

# Optional: Reset index to get the timestamp back as a column
filled_df = filled_df.reset_index().rename(columns={'index': 'ephoc_as_datatime'})
  • method='time': Interpolates based on the time difference between adjacent valid points (ideal for time series data)
  • If you need a different interpolation method (e.g., quadratic spline), replace 'time' with 'quadratic' or another valid pandas interpolation method

Output Verification

When you run this code, you'll get a DataFrame that matches your example:

  • The 11:22:37.170 timestamp will have att1=4 and att2=8 (average of the 165ms, 168ms, and 169ms samples)
  • All gaps between valid resampled points will be filled with interpolated values
  • The final timestamp (11:22:38.120) will retain its original value since it falls exactly on a 10ms interval

内容的提问来源于stack exchange,提问作者Carlo Allocca

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 03:57:14