You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

GitBash中如何拆分800MB文本文件?要求拆分文件首行以H开头

解决大文本文件按规则拆分的方案

原split -l 39000只能按固定行数拆分,无法保证新文件首行以'H'开头。可以用awk脚本实现你需要的逻辑:当累计行数达到39000后,找到首个以'H'开头的行作为新文件的起始,同时确保每个拆分文件的首行都是'H'。

方法一:使用独立awk脚本

  1. 创建名为split_by_H.awk的文件,写入以下代码:
BEGIN {
    max_lines = 39000  # 每个文件的最小行数
    file_num = 1
    output_file = sprintf("part_%03d.txt", file_num)  # 生成带三位编号的文件名
    line_count = 0
}

{
    if (line_count < max_lines) {
        print > output_file
        line_count++
    } else {
        # 行数达标后,寻找首个以H开头的行作为拆分点
        if ($0 ~ /^H/) {
            close(output_file)
            file_num++
            output_file = sprintf("part_%03d.txt", file_num)
            print > output_file
            line_count = 1
        } else {
            print > output_file
            line_count++
        }
    }
}

END {
    close(output_file)
}
  1. 在GitBash中执行命令:
awk -f split_by_H.awk myfilename

方法二:直接使用单行命令

如果不想创建脚本文件,可直接运行以下命令:

awk 'BEGIN{max=39000;fn=1;out=sprintf("part_%03d.txt",fn);lc=0} {if(lc<max){print>out;lc++}else{if(/^H/){close(out);fn++;out=sprintf("part_%03d.txt",fn);print>out;lc=1}else{print>out;lc++}}} END{close(out)}' myfilename

自定义调整说明

  • 若要修改每个文件的最小行数,修改max_lines(单行命令中为max)的值即可。
  • 文件名格式可通过调整sprintf中的字符串修改,比如"split_%d.txt"会生成split_1.txt、split_2.txt等。
  • 脚本会自动处理最后一个文件,即使剩余行数不足39000,只要首行是'H'就会正常生成。

内容的提问来源于stack exchange,提问作者donga garu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 10:01:19