You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python实现CSV文件中分散数据向上对齐合并

解决CSV列数据块向上对齐的Python实现

问题场景示例

假设你的CSV内容类似这样(多组分散的列数据块,组间可能有空行分隔):

Name,Points,Type
Alice,,
Bob,,
,100,
,200,
,,TypeA
,,TypeB

Charlie,,
Dave,,
,300,
,,TypeC

需要转换成:

Name,Points,Type
Alice,100,TypeA
Bob,200,TypeB
Charlie,300,TypeC

实现代码(用基础csv模块,适合新手)

import csv
from itertools import zip_longest

def align_csv_blocks(input_path, output_path):
    # 读取原始CSV数据
    with open(input_path, 'r', newline='', encoding='utf-8') as f:
        reader = csv.DictReader(f)
        rows = list(reader)
        fieldnames = reader.fieldnames

    # 1. 按列类型分组块:把连续的同列非空行归为一个块
    blocks = []
    current_block_type = None
    current_block = []
    for row in rows:
        # 判断当前行属于哪个列(仅一列非空的情况)
        non_empty_cols = [col for col in fieldnames if row[col].strip() != '']
        if not non_empty_cols:
            # 空行:结束当前块,开始新组的标记
            if current_block:
                blocks.append((current_block_type, current_block))
                current_block_type = None
                current_block = []
            continue
        # 仅处理单列非空的行(符合你的数据场景)
        if len(non_empty_cols) == 1:
            col_type = non_empty_cols[0]
            if col_type != current_block_type:
                # 块类型变化:保存之前的块
                if current_block:
                    blocks.append((current_block_type, current_block))
                current_block_type = col_type
                current_block = [row]
            else:
                current_block.append(row)
    # 保存最后一个块
    if current_block:
        blocks.append((current_block_type, current_block))

    # 2. 把块按组划分:每组包含Name、Points、Type三个块
    groups = []
    current_group = {}
    for block_type, block_rows in blocks:
        current_group[block_type] = block_rows
        # 当组内三个列的块都齐了,存入groups并重置
        if set(current_group.keys()) == {'Name', 'Points', 'Type'}:
            groups.append(current_group)
            current_group = {}
    # 处理最后一个未完成的组(如果有的话)
    if current_group:
        groups.append(current_group)

    # 3. 对齐每组的列数据,生成完整行
    aligned_rows = []
    for group in groups:
        # 提取每组各列的非空值列表
        name_list = [row['Name'].strip() for row in group.get('Name', []) if row['Name'].strip()]
        points_list = [row['Points'].strip() for row in group.get('Points', []) if row['Points'].strip()]
        type_list = [row['Type'].strip() for row in group.get('Type', []) if row['Type'].strip()]

        # 用zip_longest对齐,短列表补空字符串
        for name, points, type_val in zip_longest(name_list, points_list, type_list, fillvalue=''):
            aligned_rows.append({
                'Name': name,
                'Points': points,
                'Type': type_val
            })

    # 4. 写入对齐后的CSV
    with open(output_path, 'w', newline='', encoding='utf-8') as f:
        writer = csv.DictWriter(f, fieldnames=fieldnames)
        writer.writeheader()
        writer.writerows(aligned_rows)

# 使用示例:替换成你的输入输出路径
align_csv_blocks('input.csv', 'output.csv')

代码说明

  1. 块分组:遍历每行,识别连续的单列非空行,将其归为对应列的块(比如连续的Name非空行组成Name块),空行作为组的分隔。
  2. 组划分:当收集到Name、Points、Type三个块时,视为一组,开始处理下一组。
  3. 数据对齐:提取每组各列的非空值列表,用zip_longest按顺序一一对应,短列表自动补空,生成完整数据行。
  4. 写入结果:将对齐后的行写入新CSV,保留原始表头。

适配你的场景

如果你的数据组之间没有空行分隔,代码依然能自动识别块类型的切换(比如Name块结束后是Points块,再是Type块,之后又是Name块,会自动分成两组)。如果存在多列同时非空的行,代码会跳过这类行(你可以根据需求调整non_empty_cols的判断逻辑)。

内容的提问来源于stack exchange,提问作者garrettk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 03:32:49