You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何拆分带表头的超大文本文件并删除原文件?

拆分超大带表头文件并边拆边删原文件内容

方法一:Bash 脚本实现(快速高效)

步骤说明

  1. 提取原文件表头保存为临时文件
  2. 循环读取原文件表头后的内容,按1GB拆分到新文件(每个新文件先写入表头)
  3. 每完成一个拆分块,立即截断原文件删除已处理部分

脚本代码

#!/bin/bash

INPUT_FILE="large_file.txt"
CHUNK_SIZE="1G"  # 按1GB拆分,可按需调整
HEADER_TMP="header.tmp"

# 提取表头
head -n 1 "$INPUT_FILE" > "$HEADER_TMP"
# 计算表头字节数,用于判断原文件是否只剩表头
HEADER_SIZE=$(stat -c %s "$HEADER_TMP")

chunk_num=1

while true; do
    FILE_SIZE=$(stat -c %s "$INPUT_FILE")
    # 原文件只剩表头或为空时退出循环
    if [ "$FILE_SIZE" -le "$HEADER_SIZE" ]; then
        break
    fi

    # 创建拆分文件并写入表头
    cat "$HEADER_TMP" > "chunk_${chunk_num}.txt"
    # 从原文件跳过表头,读取1GB数据追加到拆分文件
    dd if="$INPUT_FILE" of="chunk_${chunk_num}.txt" bs="$CHUNK_SIZE" skip=1 seek=1 conv=notrunc oflag=append status=progress
    # 截断原文件,移除已处理的1GB数据
    truncate -s "-$CHUNK_SIZE" "$INPUT_FILE"

    chunk_num=$((chunk_num + 1))
done

# 处理最后剩余的不足1GB的部分
FILE_SIZE=$(stat -c %s "$INPUT_FILE")
if [ "$FILE_SIZE" -gt "$HEADER_SIZE" ]; then
    cat "$HEADER_TMP" > "chunk_${chunk_num}.txt"
    # 追加原文件表头后的剩余内容
    tail -n +2 "$INPUT_FILE" >> "chunk_${chunk_num}.txt"
    # 清空原文件
    > "$INPUT_FILE"
fi

# 清理临时文件
rm "$HEADER_TMP"

echo "拆分完成,共生成 $chunk_num 个文件"

方法二:Python 脚本实现(适配后续Python分析)

适合直接衔接Python数据处理流程,全程流式处理,无需加载整个文件到内存。

脚本代码

import os

def split_large_file(input_path, chunk_size_gb=1):
    CHUNK_SIZE = chunk_size_gb * 1024 * 1024 * 1024  # 转换为字节单位
    chunk_num = 1

    # 读取表头,按实际文件编码调整encoding参数(如'gbk')
    with open(input_path, 'r', encoding='utf-8') as f:
        header = f.readline()
    header_len = len(header.encode('utf-8'))

    while True:
        file_size = os.path.getsize(input_path)
        # 原文件只剩表头或为空时停止循环
        if file_size <= header_len:
            break

        # 计算本次读取的最大字节数
        read_bytes = min(CHUNK_SIZE, file_size - header_len)
        chunk_path = f"chunk_{chunk_num}.txt"

        # 写入表头+数据块
        with open(chunk_path, 'w', encoding='utf-8') as chunk_f:
            chunk_f.write(header)
            with open(input_path, 'r', encoding='utf-8') as f:
                f.seek(header_len)
                chunk_f.write(f.read(read_bytes))

        # 截断原文件,删除已处理部分
        os.truncate(input_path, file_size - read_bytes)
        chunk_num += 1

    # 处理最后剩余的小数据块
    file_size = os.path.getsize(input_path)
    if file_size > header_len:
        chunk_path = f"chunk_{chunk_num}.txt"
        with open(chunk_path, 'w', encoding='utf-8') as chunk_f:
            chunk_f.write(header)
            with open(input_path, 'r', encoding='utf-8') as f:
                f.seek(header_len)
                chunk_f.write(f.read())
        # 清空原文件
        os.truncate(input_path, 0)

    print(f"拆分完成,共生成 {chunk_num} 个文件")

if __name__ == "__main__":
    split_large_file("large_file.txt", chunk_size_gb=1)

注意事项

  • 操作前务必备份原文件!截断操作不可逆,出错会导致数据丢失
  • 若文件编码不是UTF-8,需修改脚本中的encoding参数
  • Bash脚本中dd的bs参数,部分系统需用1024M替代1G,或直接写字节数1073741824
  • Python脚本若内存不足,可将一次性读取改为分小块循环写入(比如每次读1MB)

内容的提问来源于stack exchange,提问作者resourcefulodysseus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 08:43:33