You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按header和footer分割文本文件并批量保存分段内容?

需求与问题背景

现有一个格式如下的文本文件:

[timestamp1] header with space
[timestamp2] data1 
[timestamp3] data2
[timestamp4] data3
[timestamp5] ..
[timestamp6] footer with space
[timestamp7] junk
[timestamp8] header with space
[timestamp9] data4
[timestamp10] data5
[timestamp11] ...
[timestamp12] footer with space
[timestamp13] junk
[timestamp14] header with space
[timestamp15] data6
[timestamp16] data7
[timestamp17] data8
[timestamp18] ..
[timestamp19] footer with space

需要提取每一段header和footer之间的内容,分别保存为file1、file2等文件(是否保留timestamp均可)。例如file1应包含:

data1
data2
data3
..

目前已能用sed命令提取第一段内容:

sed -n "/header/,/footer/{p;/footer/q}" file

但无法循环处理后续匹配内容,希望得到完整解决方案。


方法一:用awk一次性处理(推荐)

awk可以直接遍历文件自动分段输出,无需循环修改原文件,效率更高:

awk '/header/{flag=1; count++; next} /footer/{flag=0; next} flag{print > "file"count}' input.txt

代码说明:

  • 匹配到header时,开启内容标记flag=1,计数器count加1,跳过当前行(不输出header行)
  • 匹配到footer时,关闭内容标记flag=0,跳过当前行(不输出footer行)
  • 当flag=1时,将当前行输出到对应编号的file${count}文件中

如果需要去掉每行开头的[timestampX] 前缀,可修改为:

awk '/header/{flag=1; count++; next} /footer/{flag=0; next} flag{sub(/^\[[^]]+\] /,""); print > "file"count}' input.txt

方法二:sed结合shell循环处理

如果坚持使用sed,可以通过循环每次提取一段后删除已处理内容,直到文件中无header:

count=1
input_file="input.txt"
temp_file="temp_${input_file}"
cp "$input_file" "$temp_file"

while grep -q "header" "$temp_file"; do
    # 提取header和footer之间的内容(排除首尾行)
    sed -n "/header/,/footer/{/header/d;/footer/d;p}" "$temp_file" > "file${count}"
    # 删除已处理的header到footer段
    sed -i "/header/,/footer/d" "$temp_file"
    ((count++))
done

rm "$temp_file"

代码说明:

  • 先复制原文件到临时文件,避免修改原始数据
  • 循环判断临时文件中是否存在header,直到所有段处理完成
  • 每次提取目标内容后,删除临时文件中已处理的段落,避免重复提取

内容的提问来源于stack exchange,提问作者mehdi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 16:30:49