You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python基于DataFrame的ID值拆分CSV为多个子文件

Python实现按ID拆分CSV文件

下面是两种实用的实现方式,分别基于标准库和第三方库,按需选择:

方法一:使用Python标准库csv模块

无需额外安装依赖,适合轻量级数据处理:

import csv
from collections import defaultdict

# 原始CSV文件路径
input_file = "original_data.csv"

# 用字典存储每个ID对应的文件对象和写入器
file_handlers = defaultdict(lambda: None)
headers = None

with open(input_file, mode='r', newline='', encoding='utf-8') as infile:
    reader = csv.DictReader(infile)
    headers = reader.fieldnames  # 获取表头
    
    for row in reader:
        current_id = row['ID']
        # 首次遇到该ID时,创建对应输出文件并写入表头
        if file_handlers[current_id] is None:
            output_file = f"{current_id}.csv"
            outfile = open(output_file, mode='w', newline='', encoding='utf-8')
            writer = csv.DictWriter(outfile, fieldnames=headers)
            writer.writeheader()
            file_handlers[current_id] = (outfile, writer)
        
        # 将当前行写入对应文件
        file_handlers[current_id][1].writerow(row)

# 关闭所有输出文件
for outfile, _ in file_handlers.values():
    outfile.close()

方法二:使用pandas库

代码更简洁,适合处理较大数据集,需先安装pandas(pip install pandas):

import pandas as pd

# 读取原始CSV数据
df = pd.read_csv("original_data.csv")

# 按ID分组,逐个保存为独立CSV
for id_val, group in df.groupby('ID'):
    group.to_csv(f"{id_val}.csv", index=False)

补充说明

  • 两种方法都会保留原始CSV的表头,每个输出文件仅包含对应ID的所有行
  • 如果原始CSV采用非UTF-8编码,可修改代码中的encoding参数(比如gbk)适配

内容的提问来源于stack exchange,提问作者Akansha

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 19:40:29