You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从指定格式原始输入数据按规则构建符合要求的DataFrame

数据整理方案

原始数据

$0011:0524-08-2021
$0021:0624-08-2021
&0011:0724-08-2021
&0021:0924-08-2021
$0031:3124-08-2021
&0031:3224-08-2021
$0041:3924-08-2021
&0041:3924-08-2021
$0012:3124-08-2021
&0012:3324-08-2021

记录格式规则

  • 以$开头的为起始记录:例如$0011:0524-08-2021中,第2-4位001为ID,第5-8位1:05为时间,最后10位24-08-2021为日期
  • 以&开头的为结束记录:字段规则与起始记录完全一致,仅开头标识代表结束含义

整理要求

  1. DataFrame第1列仅存储$开头的起始记录,第2列仅存储&开头的结束记录
  2. 两列数据按记录中的时间升序排列,相同ID的起始、结束记录一一对应

实现代码(Python + Pandas)

import pandas as pd

# 导入原始数据
raw_list = [
    "$0011:0524-08-2021", "$0021:0624-08-2021", "&0011:0724-08-2021", "&0021:0924-08-2021",
    "$0031:3124-08-2021", "&0031:3224-08-2021", "$0041:3924-08-2021", "&0041:3924-08-2021",
    "$0012:3124-08-2021", "&0012:3324-08-2021"
]

# 拆分起始/结束记录
start_list = [i for i in raw_list if i.startswith("$")]
end_list = [i for i in raw_list if i.startswith("&")]

# 定义匹配键提取规则(ID+日期保证同组匹配,时间用于排序)
def get_attr(record):
    return {
        "match_key": f"{record[1:4]}_{record[-10:]}",
        "sort_time": record[4:8]
    }

# 分别转为DataFrame后合并
start_df = pd.DataFrame([{"start": i, **get_attr(i)} for i in start_list])
end_df = pd.DataFrame([{"end": i, **get_attr(i)} for i in end_list])
res_df = pd.merge(start_df, end_df, on="match_key")

# 按时间升序排序后输出结果
res_df = res_df.sort_values("sort_time").reset_index(drop=True)[["start", "end"]]
print(res_df)

输出结果

start                  end
0  $0011:0524-08-2021  &0011:0724-08-2021
1  $0021:0624-08-2021  &0021:0924-08-2021
2  $0031:3124-08-2021  &0031:3224-08-2021
3  $0012:3124-08-2021  &0012:3324-08-2021
4  $0041:3924-08-2021  &0041:3924-08-2021

内容的提问来源于stack exchange,提问作者spectre

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.07 03:15:04