如何按header和footer分割文本文件并批量保存分段内容?
需求与问题背景
现有一个格式如下的文本文件:
[timestamp1] header with space [timestamp2] data1 [timestamp3] data2 [timestamp4] data3 [timestamp5] .. [timestamp6] footer with space [timestamp7] junk [timestamp8] header with space [timestamp9] data4 [timestamp10] data5 [timestamp11] ... [timestamp12] footer with space [timestamp13] junk [timestamp14] header with space [timestamp15] data6 [timestamp16] data7 [timestamp17] data8 [timestamp18] .. [timestamp19] footer with space
需要提取每一段header和footer之间的内容,分别保存为file1、file2等文件(是否保留timestamp均可)。例如file1应包含:
data1 data2 data3 ..
目前已能用sed命令提取第一段内容:
sed -n "/header/,/footer/{p;/footer/q}" file
但无法循环处理后续匹配内容,希望得到完整解决方案。
方法一:用awk一次性处理(推荐)
awk可以直接遍历文件自动分段输出,无需循环修改原文件,效率更高:
awk '/header/{flag=1; count++; next} /footer/{flag=0; next} flag{print > "file"count}' input.txt
代码说明:
- 匹配到
header时,开启内容标记flag=1,计数器count加1,跳过当前行(不输出header行) - 匹配到
footer时,关闭内容标记flag=0,跳过当前行(不输出footer行) - 当
flag=1时,将当前行输出到对应编号的file${count}文件中
如果需要去掉每行开头的[timestampX] 前缀,可修改为:
awk '/header/{flag=1; count++; next} /footer/{flag=0; next} flag{sub(/^\[[^]]+\] /,""); print > "file"count}' input.txt
方法二:sed结合shell循环处理
如果坚持使用sed,可以通过循环每次提取一段后删除已处理内容,直到文件中无header:
count=1 input_file="input.txt" temp_file="temp_${input_file}" cp "$input_file" "$temp_file" while grep -q "header" "$temp_file"; do # 提取header和footer之间的内容(排除首尾行) sed -n "/header/,/footer/{/header/d;/footer/d;p}" "$temp_file" > "file${count}" # 删除已处理的header到footer段 sed -i "/header/,/footer/d" "$temp_file" ((count++)) done rm "$temp_file"
代码说明:
- 先复制原文件到临时文件,避免修改原始数据
- 循环判断临时文件中是否存在
header,直到所有段处理完成 - 每次提取目标内容后,删除临时文件中已处理的段落,避免重复提取
内容的提问来源于stack exchange,提问作者mehdi
相关产品推荐
相关产品推荐

