You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Pandas中对多时段数据集插值并忽略缺失值的技术需求

多时间块内的插值数据获取问题

我需要在多个不同时间块上获取插值数据,现有数据集如下:

>>> import pandas as pd
>>> test_data = pd.read_csv("test_data.csv")
>>> test_data
     ID  Condition_num Condition_type  Rating  Timestamp_ms
0   101              1        Active     58.0            30
1   101              1        Active     59.0            60
2   101              1        Active     65.0            90
3   101              1        Active     70.0           120
4   101              1        Active     80.0           150
5   101              2          Break     NaN           180
6   101              3       Active_2    55.0           210
7   101              3       Active_2    60.0           240
8   101              3       Active_2    63.0           270
9   101              3       Active_2    70.0           300
10  101              4          Break     NaN           330
11  101              5       Active_3    69.0           360
12  101              5       Active_3    71.0           390
13  101              5       Active_3    50.0           420
14  101              5       Active_3    41.0           450
15  101              5       Active_3    43.0           480

我需要将时间戳列重采样为40ms的间隔以匹配外部数据集,当前使用的处理代码如下:

# 将时间戳列设置为正确单位的datetime,然后设置为索引
test_data['Timestamp_ms'] = pd.to_datetime(test_data['Timestamp_ms'], unit='ms')
test_data = test_data.set_index('Timestamp_ms')

# 重采样索引从0开始,先重采样到最高分辨率1ms,再重采样到40ms
test_data = test_data.reindex(
    pd.date_range(start=pd.to_datetime(0, unit='ms'), end=test_data.index.max(), freq='ms')
)

test_data = test_data.resample('1ms').interpolate().resample('40ms').interpolate()

# 将Rating值四舍五入为整数
test_data.Rating = test_data.Rating.round()

得到的结果如下:

ID  Condition_num Condition_type     Rating
1970-01-01 00:00:00.000    NaN            NaN            NaN        NaN
1970-01-01 00:00:00.040  101.0       1.000000            NaN  58.333333
1970-01-01 00:00:00.080  101.0       1.000000            NaN  63.000000
1970-01-01 00:00:00.120  101.0       1.000000        Active   70.000000
1970-01-01 00:00:00.160  101.0       1.333333            NaN  75.833333
1970-01-01 00:00:00.200  101.0       2.666667            NaN  59.166667
1970-01-01 00:00:00.240  101.0       3.000000       Active_2  60.000000
1970-01-01 00:00:00.280  101.0       3.000000            NaN  65.333333
1970-01-01 00:00:00.320  101.0       3.666667            NaN  69.666667
1970-01-01 00:00:00.360  101.0       5.000000       Active_3  69.000000
1970-01-01 00:00:00.400  101.0       5.000000            NaN  64.000000
1970-01-01 00:00:00.440  101.0       5.000000            NaN  44.000000
1970-01-01 00:00:00.480  101.0       5.000000       Active_3  43.000000

当前问题与需求

目前的处理方式存在两个核心问题:

  • 无法区分哪些Rating属于Active时段
  • 无法判断当前Rating是否来自Break时段的外推

核心需求:仅在Active块内进行插值,同时保持所有数据对齐整个数据集的起始时间。

我曾尝试将NaN的Rating设为0并在每个条件开头进行插值,但这反而更严重地改变了Rating值,希望得到可行的解决方案。


内容的提问来源于stack exchange,提问作者Stephen J. Suss Chacán

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 09:51:27