You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas将堆叠格式CSV文件转换为标准结构DataFrame?

使用Pandas转换堆叠式CSV为标准DataFrame

解决方案步骤

  1. 读取CSV文件并按空行分割数据块
  2. 逐个处理每个信号数据块,提取信号名称和时间序列数据
  3. 将每个信号的DataFrame按时间戳合并为宽格式

完整代码

import pandas as pd

# 读取CSV文件,过滤空行
with open('input.csv', 'r', encoding='utf-8') as f:
    lines = [line.strip() for line in f if line.strip()]

# 分割为独立的信号数据块
blocks = []
current_block = []
for line in lines:
    if line.startswith('Trace Name,'):
        if current_block:
            blocks.append(current_block)
        current_block = [line]
    else:
        current_block.append(line)
if current_block:
    blocks.append(current_block)

# 处理每个数据块,生成单个信号的DataFrame
signal_dfs = []
for block in blocks:
    # 提取信号名称
    signal_name = block[0].split(',')[1]
    # 提取时间戳和数值数据(跳过前3行的描述信息)
    data_rows = [row.split(',', 1) for row in block[3:]]
    temp_df = pd.DataFrame(data_rows, columns=['Timestamp', signal_name])
    
    # 转换时间戳为datetime类型(处理EDT时区)
    temp_df['Timestamp'] = pd.to_datetime(temp_df['Timestamp'])
    # 转换数值为数字类型
    temp_df[signal_name] = pd.to_numeric(temp_df[signal_name])
    
    signal_dfs.append(temp_df)

# 合并所有信号的DataFrame,按时间戳对齐
final_df = signal_dfs[0]
for df in signal_dfs[1:]:
    final_df = pd.merge(final_df, df, on='Timestamp', how='outer')

# 可选:格式化时间戳为示例中的显示格式(去除时区和毫秒)
final_df['Timestamp'] = final_df['Timestamp'].dt.strftime('%m/%d/%Y %H:%M:%S')

# 重置索引
final_df = final_df.reset_index(drop=True)

print(final_df)

代码说明

  • 读取与分割: 先读取所有非空行,再按Trace Name,开头的行分割成独立的信号块,确保每个块对应一个信号的数据。
  • 数据提取: 每个块的第一行提取信号名称,从第四行开始是实际的时间戳-数值对,转换为临时DataFrame。
  • 类型转换: 将时间戳解析为datetime对象(自动处理EDT时区),数值转换为数字类型避免字符串问题。
  • 合并对齐: 使用merge按时间戳合并所有信号DataFrame,how='outer'确保即使时间戳不完全一致也能保留所有数据,缺失值会填充为NaN。
  • 格式调整: 可选步骤将时间戳格式化为示例中的短格式,根据实际需求调整即可。

内容的提问来源于stack exchange,提问作者Hannibal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 17:07:19