You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

从Azure Blob分块下载文件后合并为单个CSV的方法

解决Azure Storage分块下载CSV后的不完整行合并问题

优化方案一:直接流式下载到单个文件(推荐)

你之前的代码把每个字节块存成独立文件,这才导致了行被切割的问题。其实可以直接通过流式下载,将分块的字节依次写入同一个目标文件,最终得到和原文件完全一致的完整CSV,无需后续合并操作。

import os
from azure.storage.blob import BlobClient

def download_large_csv(blob_client, dest_file_path):
    chunk_size = 100 * 1024 * 1024  # 100 MB 分块大小
    with open(dest_file_path, "wb") as dest_file:
        # 启用流式下载,设置并发数提升下载速度
        stream = blob_client.download_blob(max_concurrency=4)
        # 按分块读取并写入目标文件
        for chunk in stream.chunks(chunk_size=chunk_size):
            dest_file.write(chunk)

使用这个方法,下载过程中不会生成多个零散的chunk文件,直接得到完整的CSV,彻底避免行切割问题。

方案二:合并已下载的chunk文件(针对已存在的分块文件)

如果已经下载了多个chunk文件,可以通过以下步骤合并,修复不完整行并处理重复表头:

import os

def merge_csv_chunks(chunk_folder, dest_file_path):
    # 按序号排序chunk文件,确保顺序正确
    chunk_files = sorted(
        [f for f in os.listdir(chunk_folder) if f.startswith("chunk_") and f.endswith(".csv")],
        key=lambda x: int(x.split("_")[1].split(".")[0])
    )
    
    with open(dest_file_path, "wb") as dest_file:
        leftover = b""  # 存储上一个chunk末尾的不完整字节段
        
        for idx, chunk_file_name in enumerate(chunk_files):
            chunk_path = os.path.join(chunk_folder, chunk_file_name)
            with open(chunk_path, "rb") as chunk_file:
                chunk_data = chunk_file.read()
                # 拼接上一个chunk的剩余内容
                full_data = leftover + chunk_data
                # 找到最后一个换行符的位置,分割完整行和不完整行
                last_newline = full_data.rfind(b"\n")
                
                if last_newline != -1:
                    # 写入所有完整行
                    dest_file.write(full_data[:last_newline+1])
                    # 保存剩余的不完整部分,用于和下一个chunk拼接
                    leftover = full_data[last_newline+1:]
                else:
                    # 当前chunk无完整换行,全部作为剩余内容
                    leftover = full_data
            
            # 处理最后一个chunk的剩余内容
            if idx == len(chunk_files) - 1 and leftover:
                dest_file.write(leftover)
    
    # 清理重复表头(如果原CSV包含表头)
    with open(dest_file_path, "r", encoding="utf-8") as f:
        lines = f.readlines()
    
    # 检查是否存在重复表头行
    if len(lines) > 1 and lines[0].strip() == lines[1].strip():
        cleaned_lines = [lines[0]] + [line for line in lines[1:] if line.strip() != lines[0].strip()]
    else:
        cleaned_lines = lines
    
    with open(dest_file_path, "w", encoding="utf-8") as f:
        f.writelines(cleaned_lines)

注意事项

  • 编码适配:如果你的CSV使用非UTF-8编码(如GBK),请修改代码中的encoding参数。
  • 换行符兼容:如果CSV使用Windows风格的\r\n换行,代码中的b"\n"仍能正常工作,无需修改;若需严格匹配,可替换为b"\r\n"。
  • 无表头场景:如果原CSV没有表头,可删除代码中清理重复表头的部分。

内容的提问来源于stack exchange,提问作者Winston Li

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 13:07:49