You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

按第三列实体拆分CSV文件(各实体列数不固定)

拆分多列数实体的CSV文件解决方案

直接用Python写个脚本就能搞定,核心思路是按第三列的实体标识分组,再根据每个实体的列数截取内容后导出独立文件:

import csv
from collections import defaultdict

# 存储每个实体对应的行数据(包括后续要处理的表头)
entity_data = defaultdict(list)
# 记录每个实体对应的列数
entity_col_counts = {}

# 读取合并后的CSV文件
with open('combine.csv', 'r', newline='', encoding='utf-8') as infile:
    reader = csv.reader(infile)
    # 先读取表头
    header = next(reader)
    # 遍历所有数据行
    for row in reader:
        if len(row) < 3:
            continue  # 跳过没有实体标识的无效行
        entity_id = row[2]
        # 用该实体的第一行长度作为列数标准
        if entity_id not in entity_col_counts:
            entity_col_counts[entity_id] = len(row)
        # 将当前行加入对应实体的数据集
        entity_data[entity_id].append(row)

# 为每个实体生成独立CSV文件
for entity, rows in entity_data.items():
    col_count = entity_col_counts[entity]
    # 截取对应列数的表头
    entity_header = header[:col_count]
    with open(f'{entity}.csv', 'w', newline='', encoding='utf-8') as outfile:
        writer = csv.writer(outfile)
        writer.writerow(entity_header)
        # 每行只保留对应列数的内容再写入
        for row in rows:
            writer.writerow(row[:col_count])

关键说明

  • 脚本会自动识别每个实体的列数(以该实体第一行的列数为标准)
  • 自动过滤列数不足3的无效行(避免找不到实体标识)
  • 导出的文件直接以实体命名,表头和数据都会按对应实体的列数截取

内容的提问来源于stack exchange,提问作者hardik rawal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 09:22:40