You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python复杂CSV解析失败,求实现指定层级数据提取逻辑

解析特殊结构CSV为层级字典的实现方案

问题概述

需要解析一份非标准结构的CSV文件:文件通过Start或Start repeat标记操作起始,后续行依次包含Read/Write操作类型、地址信息,以及多组Data数据,最终要将这些内容组织成指定的嵌套字典结构。现有基于pandas的实现无法正确关联各部分内容,尤其是多组数据的收集。

样本CSV内容

,,Start,,
,Read,0x1000,,
,Data,0x01,0x02,0x03
,,Start repeat,,
,Write,0x2000,,
,Data,0x04,0x05
,Data,0x06,0x07,0x08

期望输出结构

{
    "operations": [
        {
            "type": "Read",
            "address": "0x1000",
            "data": [["0x01", "0x02", "0x03"]]
        },
        {
            "type": "Write",
            "address": "0x2000",
            "data": [["0x04", "0x05"], ["0x06", "0x07", "0x08"]]
        }
    ]
}

现有pandas尝试代码(存在问题)

import pandas as pd

df = pd.read_csv("special.csv", header=None)
result = {}
operations = []

for idx, row in df.iterrows():
    if pd.notna(row[2]) and (row[2] == "Start" or row[2] == "Start repeat"):
        op = {}
        next_row = df.iloc[idx+1]
        op["type"] = next_row[1]
        op["address"] = next_row[2]
        # 无法正确收集多组Data行,逻辑缺失
        operations.append(op)

result["operations"] = operations
print(result)

解决方案:逐行解析替代pandas

由于该CSV结构不符合标准表格格式,pandas的表格化处理会增加复杂度,直接逐行读取解析更灵活可靠。

完整实现代码

def parse_special_csv(file_path):
    result = {"operations": []}
    current_operation = None

    with open(file_path, 'r') as csv_file:
        for line in csv_file:
            # 清洗行数据:分割、去空白、过滤空元素
            cleaned_parts = [segment.strip() for segment in line.strip().split(',') if segment.strip()]
            if not cleaned_parts:
                continue

            # 触发新操作的起始标记
            if cleaned_parts[0] in ("Start", "Start repeat"):
                # 若存在未完成的操作,先加入结果(处理异常行情况)
                if current_operation:
                    result["operations"].append(current_operation)
                # 初始化新操作的字典
                current_operation = {"type": None, "address": None, "data": []}
            
            # 匹配Read/Write操作类型与地址
            elif cleaned_parts[0] in ("Read", "Write") and current_operation:
                current_operation["type"] = cleaned_parts[0]
                current_operation["address"] = cleaned_parts[1]
            
            # 收集Data行的数据组
            elif cleaned_parts[0] == "Data" and current_operation:
                current_operation["data"].append(cleaned_parts[1:])

    # 处理最后一个未加入结果的操作
    if current_operation:
        result["operations"].append(current_operation)
    
    return result

# 测试示例
if __name__ == "__main__":
    import json
    parsed_data = parse_special_csv("special.csv")
    print(json.dumps(parsed_data, indent=4))

代码说明

  • 逐行处理:避开pandas对非标准表格的解析限制,直接读取每行内容
  • 状态跟踪:用current_operation变量跟踪当前正在构建的操作,关联起始标记、操作类型、地址和数据
  • 数据清洗:过滤每行中的空字符串和空白内容,只保留有效数据段
  • 分支处理:分别识别三种核心行类型,逐步填充操作字典,最后统一收集到结果中

验证结果

运行上述代码后,输出将完全匹配期望的嵌套字典结构,能够正确处理Start/Start repeat两种起始标记,以及多组Data行的收集。

内容的提问来源于stack exchange,提问作者user2625119

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 01:45:04