按第三列实体拆分CSV文件(各实体列数不固定)
拆分多列数实体的CSV文件解决方案
直接用Python写个脚本就能搞定,核心思路是按第三列的实体标识分组,再根据每个实体的列数截取内容后导出独立文件:
import csv from collections import defaultdict # 存储每个实体对应的行数据(包括后续要处理的表头) entity_data = defaultdict(list) # 记录每个实体对应的列数 entity_col_counts = {} # 读取合并后的CSV文件 with open('combine.csv', 'r', newline='', encoding='utf-8') as infile: reader = csv.reader(infile) # 先读取表头 header = next(reader) # 遍历所有数据行 for row in reader: if len(row) < 3: continue # 跳过没有实体标识的无效行 entity_id = row[2] # 用该实体的第一行长度作为列数标准 if entity_id not in entity_col_counts: entity_col_counts[entity_id] = len(row) # 将当前行加入对应实体的数据集 entity_data[entity_id].append(row) # 为每个实体生成独立CSV文件 for entity, rows in entity_data.items(): col_count = entity_col_counts[entity] # 截取对应列数的表头 entity_header = header[:col_count] with open(f'{entity}.csv', 'w', newline='', encoding='utf-8') as outfile: writer = csv.writer(outfile) writer.writerow(entity_header) # 每行只保留对应列数的内容再写入 for row in rows: writer.writerow(row[:col_count])
关键说明
- 脚本会自动识别每个实体的列数(以该实体第一行的列数为标准)
- 自动过滤列数不足3的无效行(避免找不到实体标识)
- 导出的文件直接以实体命名,表头和数据都会按对应实体的列数截取
内容的提问来源于stack exchange,提问作者hardik rawal
相关产品推荐
相关产品推荐

