You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中按首行时间信息排序结构化数据集?

问题描述

有一组以空行分隔的结构化数据块,每个数据块的首行包含年份、月份、日期、小时、分钟等时间信息,部分数据块中还包含标准格式的时间ID(如ID:20060103150113)。需要按时间对这些数据块进行排序,得到指定顺序的结果。

Python实现方案

步骤1:读取并分割数据块

首先将原始文本按空行分割为独立的数据块,过滤掉无效的空块:

raw_data = """ 2006  1 3 1501 12.0 L  28.024  56.808 14.1  INS  3 0.0 3.9LINS                1
 GAP=158        0.32       2.7     8.0 13.4 -0.2096E+02 -0.1022E+03  0.3316E+02E
 ACTION:UPD 06-08-20 20:40 OP:moh  STATUS:               ID:20060103150113 L   I
 2006-01-03-1500-10S.NSN___036                                                 6
 STAT SP IPHASW D HRMM SECON CODA AMPLIT PERI AZIMU VELO AIN AR TRES W  DIS CAZ7
 BNDS BZ EPG      15 1 27.99                              90     0.010 93.4 222 
 BNDS BN ESG      15 1 39.77                              90     0.010 93.4 222 
 BNDS BE  AML     15 1 43.17      1880.4 0.39                          93.4 222 
 BNDS BN  AML     15 1 48.04      2173.7 0.56                          93.4 222 
 KRBR BZ EPG      15 1 47.76                              90     0.010  217 359 
 KRBR BN ESG      15 2 13.87                              90     0.010  217 359 
 KRBR BE  AML     15 2 17.46      2411.2 0.45                           217 359 
 KRBR BN  AML     15 2 17.52      2481.1 0.64                           217 359 
 ZHSF BZ EPN      15 2 12.69                              52     0.0 9  425  65 

 2005  110 1521 29.5 L  26.988  55.807 14.1  INS  3 0.1 3.2LINS                1
 GAP=260        0.41       6.4     1.9  5.5 -0.7704E+01 -0.3485E+01  0.2469E+02E
 ACTION:UPD 06-08-20 20:40 OP:moh  STATUS:               ID:20060110151405 L   I
 2006-01-10-1514-05S.NSN___039                                                 6
 STAT SP IPHASW D HRMM SECON CODA AMPLIT PERI AZIMU VELO AIN AR TRES W  DIS CAZ7
 BNDS BZ EPG      1521 39.93                              90     0.010 58.1  38 
 BNDS BE ESG      1521 47.41                              90    -0.110 58.1  38 
 BNDS BE  AML     1522 16.19      1554.5 0.64                          58.1  38 
 BNDS BN  AML     1522 17.86      1657.6 0.60                          58.1  38 
 GHIR BZ EPN      1522 16.05                              52    -0.110  313 298 
 GHIR BN ESG      1522 57.55                              90     0.010  313 298 
 KRBR BZ EPN      1522 20.37                              52     0.1 9  345  15 
 KRBR BE  AML     1523 16.05        72.3 0.64                           345  15 
 KRBR BN  AML     1523 27.24        63.0 0.52                           345  15 
"""
# 分割数据块并过滤空内容
blocks = [block.strip() for block in raw_data.split('\n\n') if block.strip()]

步骤2:提取数据块的时间戳

优先从数据块中的ID字段提取标准时间字符串(格式为YYYYMMDDHHMMSS),解析为可排序的datetime对象;若未找到ID字段,则尝试从首行解析时间信息:

import re
from datetime import datetime

def get_timestamp(block):
    # 优先从ID行提取标准时间
    id_match = re.search(r'ID:(\d{14})', block)
    if id_match:
        time_str = id_match.group(1)
        return datetime.strptime(time_str, '%Y%m%d%H%M%S')
    
    # 从首行解析时间(兼容两种格式)
    first_line = block.split('\n')[0].strip()
    parts = first_line.split()
    if len(parts) >= 4:
        year = int(parts[0])
        # 处理月日组合(如110表示1月10日)或分离的月日
        if len(parts[1]) == 3:
            month = int(parts[1][0])
            day = int(parts[1][1:])
            hour_min = parts[2]
            second = float(parts[3])
        else:
            month = int(parts[1])
            day = int(parts[2])
            hour_min = parts[3]
            second = float(parts[4])
        # 解析时分(如1501表示15点01分)
        hour = int(hour_min[:2])
        minute = int(hour_min[2:])
        return datetime(year, month, day, hour, minute, int(second), int((second - int(second))*1e6))
    
    # 解析失败时返回最早时间,确保该块排在最后
    return datetime.min

步骤3:按时间排序并输出

使用sorted函数以时间戳为键排序数据块,按需求选择升序或降序,最后用空行连接输出:

# 按时间降序排序(与期望结果一致)
sorted_blocks = sorted(blocks, key=get_timestamp, reverse=True)

# 输出排序后的结果
print('\n\n'.join(sorted_blocks))

代码说明

  • 数据分割:通过split('\n\n')实现空行分隔,strip()去除块前后的多余空白,过滤空块避免无效处理。
  • 时间提取:优先使用ID字段的标准时间,保证解析准确性;首行解析兼容两种格式(月日分离/合并),覆盖不同数据情况。
  • 排序逻辑:利用datetime对象的可比较特性,通过sorted函数实现高效排序,reverse=True控制降序(若需升序则改为False)。

内容的提问来源于stack exchange,提问作者Hossein Kianimehr

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 17:00:53