如何用Python解析含不等维度子表的CSV文件并拆分?
解决CSV文件拆分时的KeyError问题
问题背景
我有一个非标准格式的CSV文件,样本数据如下:
[Network] Network Settings RECORDNAME,DATA UTDFVERSION,8 Metric,0 yellowTime,3.5 allRedTime,1.0 Walk,7.0 DontWalk,11.0 HV,0.02 PHF,0.92 [Nodes] Node Data INTID,TYPE,X,Y,Z,DESCRIPTION,CBD,Inside Radius,Outside Radius,Roundabout Lanes,Circle Speed 1,1,111152,12379,0,,,,,, 2,1,134346,12311,0,,,,,, 3,3,133315,12317,0,,,,,, 4,1,133284,13574,0,,,,,,
需要用Python将其拆分为两个独立表格,但编写的代码在if row['RECORDNAME'] == 'Network Settings':处抛出KeyError,代码如下:
# Open the file with open('filename.csv', 'r') as f: # Create a reader reader = csv.DictReader(f) # Initialize empty lists for the tables network_table = [] nodes_table = [] # Loop through the rows for row in reader: # Check if the row contains the "RECORDNAME" key if 'RECORDNAME' in row: # Check if the row belongs to the "Network" or "Nodes" section if row['RECORDNAME'] == 'Network Settings': # Add the row to the "Network" table network_table.append(row) elif row['INTID'] is not None: # Add the row to the "Nodes" table nodes_table.append(row) # Print the tables print(network_table) print(nodes_table)
错误原因
这个文件不是标准CSV格式,它包含分节标记([Network]、[Nodes])、节标题行(Network Settings、Node Data),之后才是各节的表头和数据。而csv.DictReader会默认把文件第一行([Network])当作整个表格的表头,导致后续所有行的键都不符合预期——根本不存在RECORDNAME这个键,自然抛出KeyError。
解决方案
需要逐行读取文件,手动识别节、表头和数据行,代码如下:
network_table = [] nodes_table = [] with open('filename.csv', 'r') as f: # 过滤空行并去除每行首尾空格 lines = [line.strip() for line in f if line.strip()] current_section = None header = None for line in lines: # 识别节标记,切换当前处理的节 if line.startswith('[') and line.endswith(']'): current_section = line.strip('[]') header = None # 切换节后重置表头 continue # 处理Network节的表头 if current_section == 'Network' and header is None: if line == 'Network Settings': continue # 跳过节标题行 header = line.split(',') continue # 处理Nodes节的表头 if current_section == 'Nodes' and header is None: if line == 'Node Data': continue # 跳过节标题行 header = line.split(',') continue # 处理数据行,将表头与值配对成字典 if current_section and header: values = line.split(',') row = dict(zip(header, values)) if current_section == 'Network': network_table.append(row) elif current_section == 'Nodes': nodes_table.append(row) # 输出结果 print("网络配置表:") for row in network_table: print(row) print("\n节点表:") for row in nodes_table: print(row)
代码说明
- 先过滤掉空行并清理每行的首尾空格,避免无效行干扰
- 通过识别
[Network]和[Nodes]标记切换处理的节 - 跳过每个节的标题行,读取下一行作为当前节的表头
- 将数据行按逗号拆分,与表头配对成字典,加入对应列表
- 自动处理数据行中的空值(比如Nodes行末尾的多个逗号会对应为空字符串)
内容的提问来源于stack exchange,提问作者joshuah9
相关产品推荐
相关产品推荐

