Python复杂CSV解析失败,求实现指定层级数据提取逻辑
解析特殊结构CSV为层级字典的实现方案
问题概述
需要解析一份非标准结构的CSV文件:文件通过Start或Start repeat标记操作起始,后续行依次包含Read/Write操作类型、地址信息,以及多组Data数据,最终要将这些内容组织成指定的嵌套字典结构。现有基于pandas的实现无法正确关联各部分内容,尤其是多组数据的收集。
样本CSV内容
,,Start,, ,Read,0x1000,, ,Data,0x01,0x02,0x03 ,,Start repeat,, ,Write,0x2000,, ,Data,0x04,0x05 ,Data,0x06,0x07,0x08
期望输出结构
{ "operations": [ { "type": "Read", "address": "0x1000", "data": [["0x01", "0x02", "0x03"]] }, { "type": "Write", "address": "0x2000", "data": [["0x04", "0x05"], ["0x06", "0x07", "0x08"]] } ] }
现有pandas尝试代码(存在问题)
import pandas as pd df = pd.read_csv("special.csv", header=None) result = {} operations = [] for idx, row in df.iterrows(): if pd.notna(row[2]) and (row[2] == "Start" or row[2] == "Start repeat"): op = {} next_row = df.iloc[idx+1] op["type"] = next_row[1] op["address"] = next_row[2] # 无法正确收集多组Data行,逻辑缺失 operations.append(op) result["operations"] = operations print(result)
解决方案:逐行解析替代pandas
由于该CSV结构不符合标准表格格式,pandas的表格化处理会增加复杂度,直接逐行读取解析更灵活可靠。
完整实现代码
def parse_special_csv(file_path): result = {"operations": []} current_operation = None with open(file_path, 'r') as csv_file: for line in csv_file: # 清洗行数据:分割、去空白、过滤空元素 cleaned_parts = [segment.strip() for segment in line.strip().split(',') if segment.strip()] if not cleaned_parts: continue # 触发新操作的起始标记 if cleaned_parts[0] in ("Start", "Start repeat"): # 若存在未完成的操作,先加入结果(处理异常行情况) if current_operation: result["operations"].append(current_operation) # 初始化新操作的字典 current_operation = {"type": None, "address": None, "data": []} # 匹配Read/Write操作类型与地址 elif cleaned_parts[0] in ("Read", "Write") and current_operation: current_operation["type"] = cleaned_parts[0] current_operation["address"] = cleaned_parts[1] # 收集Data行的数据组 elif cleaned_parts[0] == "Data" and current_operation: current_operation["data"].append(cleaned_parts[1:]) # 处理最后一个未加入结果的操作 if current_operation: result["operations"].append(current_operation) return result # 测试示例 if __name__ == "__main__": import json parsed_data = parse_special_csv("special.csv") print(json.dumps(parsed_data, indent=4))
代码说明
- 逐行处理:避开pandas对非标准表格的解析限制,直接读取每行内容
- 状态跟踪:用
current_operation变量跟踪当前正在构建的操作,关联起始标记、操作类型、地址和数据 - 数据清洗:过滤每行中的空字符串和空白内容,只保留有效数据段
- 分支处理:分别识别三种核心行类型,逐步填充操作字典,最后统一收集到结果中
验证结果
运行上述代码后,输出将完全匹配期望的嵌套字典结构,能够正确处理Start/Start repeat两种起始标记,以及多组Data行的收集。
内容的提问来源于stack exchange,提问作者user2625119
相关产品推荐
相关产品推荐

