You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python合并时间序列同时间多行并补全NOTSAMPLED值

处理异步采集的传感器时间序列数据

问题背景

我们有一份ASCII格式的时间序列数据,不同传感器的测量值通过异步采集后拼接到同一文件,数据以空格分隔。需要实现两个核心处理逻辑:

  1. 将同一时间实例的多行传感器数据合并为单行
  2. 用前序有效传感器值替换NOTSAMPLED标记

两种数据场景示例

场景一:同一时间点的传感器值分散在多行

原始数据:

2022 281 08 48 14 876 10                1.00       NOTSAMPLED          NOTSAMPLED
2022 281 08 48 14 876 10          NOTSAMPLED             0.00          NOTSAMPLED
2022 281 08 48 14 876 10          NOTSAMPLED       NOTSAMPLED                1.00
2022 281 08 48 15 391 11                1.00       NOTSAMPLED          NOTSAMPLED
2022 281 08 48 15 391 11          NOTSAMPLED             0.00          NOTSAMPLED
2022 281 08 48 15 391 11          NOTSAMPLED       NOTSAMPLED                1.00
2022 281 08 48 15 896 12                1.00       NOTSAMPLED          NOTSAMPLED
2022 281 08 48 15 896 12          NOTSAMPLED             0.00          NOTSAMPLED
2022 281 08 48 15 896 12          NOTSAMPLED       NOTSAMPLED                1.00

预期输出:

2022 281 08 48 14 876 10               1.00       0.00     1.00
2022 281 08 48 15 391 11               1.00       0.00     1.00
2022 281 08 48 15 896 12               1.00       0.00     1.00

场景二:部分时间点仅采集单个传感器值

原始数据:

2022 281 08 48 14 876 10                1.00       NOTSAMPLED          NOTSAMPLED
2022 281 08 48 14 876 10          NOTSAMPLED             0.00          NOTSAMPLED
2022 281 08 48 14 880 10          NOTSAMPLED       NOTSAMPLED               10.00
2022 281 08 48 15 391 11                1.00       NOTSAMPLED          NOTSAMPLED
2022 281 08 48 15 391 11          NOTSAMPLED             0.00          NOTSAMPLED
2022 281 08 48 15 395 11          NOTSAMPLED       NOTSAMPLED               11.00
2022 281 08 48 15 896 12                1.00       NOTSAMPLED          NOTSAMPLED
2022 281 08 48 15 896 12          NOTSAMPLED             0.00          NOTSAMPLED
2022 281 08 48 15 900 12          NOTSAMPLED       NOTSAMPLED               12.00

预期输出:

2022 281 08 48 14 876 10                1.00             0.00          NOTSAMPLED
2022 281 08 48 14 880 10                1.00             0.00               10.00
2022 281 08 48 15 391 11                1.00             0.00               10.00
2022 281 08 48 15 395 11                1.00             0.00               11.00
2022 281 08 48 15 896 12                1.00             0.00               11.00
2022 281 08 48 15 900 12                1.00             0.00               12.00

Python核心库实现方案

仅使用Python标准库完成处理,核心步骤:

  1. 按行解析数据,将同一时间戳的多行传感器值合并,保留有效数值
  2. 按时间顺序排序合并后的条目
  3. 遍历序列,用前序有效数值填充当前的NOTSAMPLED

完整代码

def process_sensor_data(input_lines):
    # 合并同一时间戳的多行数据
    time_groups = {}
    for line in input_lines:
        line = line.strip()
        if not line:
            continue
        parts = line.split()
        # 前7个字段为时间戳(年 年积日 时 分 秒 毫秒 序列号)
        time_key = tuple(parts[:7])
        sensor_values = parts[7:]
        
        if time_key not in time_groups:
            time_groups[time_key] = ['NOTSAMPLED'] * len(sensor_values)
        
        # 更新当前时间戳的传感器值,保留有效数值
        for idx, val in enumerate(sensor_values):
            if val != 'NOTSAMPLED':
                time_groups[time_key][idx] = val
    
    # 按时间顺序排序合并后的条目
    sorted_times = sorted(time_groups.keys())
    merged_data = [(time, time_groups[time]) for time in sorted_times]
    
    # 用前序有效数值填充NOTSAMPLED
    prev_valid = ['NOTSAMPLED'] * len(merged_data[0][1]) if merged_data else []
    processed = []
    
    for time_vals in merged_data:
        time, current_vals = time_vals
        new_vals = []
        for idx, val in enumerate(current_vals):
            if val == 'NOTSAMPLED':
                # 使用前序有效数值,无前序则保留NOTSAMPLED
                new_vals.append(prev_valid[idx])
            else:
                new_vals.append(val)
                # 更新前序有效数值记录
                prev_valid[idx] = val
        # 格式化输出行,保持对齐
        formatted_vals = []
        for v in new_vals:
            if v == 'NOTSAMPLED':
                formatted_vals.append(f'{v:>16}')
            else:
                formatted_vals.append(f'{float(v):>16.2f}')
        processed_line = ' '.join(time) + ' ' + ' '.join(formatted_vals)
        processed.append(processed_line)
    
    return processed

# 测试使用
if __name__ == '__main__':
    # 实际使用时可替换为读取文件的代码:
    # with open('sensor_data.txt', 'r') as f:
    #     input_lines = f.readlines()
    
    # 场景一测试
    print("场景一处理结果:")
    scenario1_input = [
        "2022 281 08 48 14 876 10                1.00       NOTSAMPLED          NOTSAMPLED",
        "2022 281 08 48 14 876 10          NOTSAMPLED             0.00          NOTSAMPLED",
        "2022 281 08 48 14 876 10          NOTSAMPLED       NOTSAMPLED                1.00",
        "2022 281 08 48 15 391 11                1.00       NOTSAMPLED          NOTSAMPLED",
        "2022 281 08 48 15 391 11          NOTSAMPLED             0.00          NOTSAMPLED",
        "2022 281 08 48 15 391 11          NOTSAMPLED       NOTSAMPLED                1.00",
        "2022 281 08 48 15 896 12                1.00       NOTSAMPLED          NOTSAMPLED",
        "2022 281 08 48 15 896 12          NOTSAMPLED             0.00          NOTSAMPLED",
        "2022 281 08 48 15 896 12          NOTSAMPLED       NOTSAMPLED                1.00"
    ]
    for line in process_sensor_data(scenario1_input):
        print(line)
    
    # 场景二测试
    print("\n场景二处理结果:")
    scenario2_input = [
        "2022 281 08 48 14 876 10                1.00       NOTSAMPLED          NOTSAMPLED",
        "2022 281 08 48 14 876 10          NOTSAMPLED             0.00          NOTSAMPLED",
        "2022 281 08 48 14 880 10          NOTSAMPLED       NOTSAMPLED               10.00",
        "2022 281 08 48 15 391 11                1.00       NOTSAMPLED          NOTSAMPLED",
        "2022 281 08 48 15 391 11          NOTSAMPLED             0.00          NOTSAMPLED",
        "2022 281 08 48 15 395 11          NOTSAMPLED       NOTSAMPLED               11.00",
        "2022 281 08 48 15 896 12                1.00       NOTSAMPLED          NOTSAMPLED",
        "2022 281 08 48 15 896 12          NOTSAMPLED             0.00          NOTSAMPLED",
        "2022 281 08 48 15 900 12          NOTSAMPLED       NOTSAMPLED               12.00"
    ]
    for line in process_sensor_data(scenario2_input):
        print(line)

代码说明

  • 时间戳合并:用元组作为时间键,将同一时间戳的多行传感器值合并,优先保留非NOTSAMPLED的有效数值
  • 时间排序:确保处理后的序列严格按时间顺序排列
  • 前序值填充:维护一个记录每个传感器最新有效数值的列表,遇到NOTSAMPLED时自动替换为前序有效数值
  • 格式对齐:输出时保持与原始数据一致的字段对齐格式,数值统一保留两位小数

内容的提问来源于stack exchange,提问作者Soumajit

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 14:45:40