文本处理:如何逆序提取Pattern A至首个Pattern B匹配区间行
Let's break down practical solutions for both of your text extraction tasks using standard Unix tools (awk, tac) — they're widely available and ideal for these jobs.
Problem 1: Extract from Pattern A to the First Pattern B
If you need to grab all lines starting from the first match of Pattern A up to (and including) the first subsequent Pattern B, here's a straightforward awk one-liner:
awk '/Pattern A/{flag=1} flag; /Pattern B/{flag=0}' example_file.txt
How this works:
- When awk hits a line matching
Pattern A, it sets aflagto 1 (active). - The
flag;line tells awk to print the current line only if the flag is on. - Once we encounter a line with
Pattern B, the flag is turned off, stopping further printing.
Edge Case Adjustments:
- No Pattern B after A: This will print everything from Pattern A to the end of the file.
- Only want the first occurrence: Add
exitafter turning off the flag to stop processing once the first Pattern B is found:awk '/Pattern A/{flag=1} flag; /Pattern B/{flag=0; exit}' example_file.txt
Problem 2: Bottom-Up Interval Extraction (AK5*R to AK2)
For this task, we need to scan from the end of the file upwards, find every line matching AK5*R, and extract the lines from that line up to the first preceding AK2. We'll name these intervals E1 (last occurrence), E2 (second-last), etc.
Full Step-by-Step Solution:
First, reverse the file so we can process from bottom to top using
tac:tac example_file.txt > reversed_file.txtUse awk to split the reversed file into target blocks and save them to numbered files:
awk ' /AK5\*R/ { if (block) { print block > "E" cnt; cnt++ } block = $0 "\n" next } /AK2/ { block = block $0 "\n" print block > "E" cnt cnt++ block = "" next } { if (block) block = block $0 "\n" } END { if (block) print block > "E" cnt } ' reversed_file.txtReverse each E file to restore the original line order:
for file in E*; do tac "$file" > "$file.tmp" && mv "$file.tmp" "$file"; done
How this works:
tacreverses the input file, so the last line of the original becomes the first line we process.- The awk script collects lines into a
blockwhen it hitsAK5*R. It keeps adding lines until it findsAK2, then saves the block to a numbered E file. If there's an unfinished block at the end (no AK2 after the last AK5*R), it saves that too. - Finally, we reverse each E file to get the lines back in their original order (since we processed reversed input).
Result:
E1contains the last occurrence ofAK5*Rin the original file, along with all lines up to the firstAK2above it.E2contains the second-lastAK5*Rinterval, and so on.
内容的提问来源于stack exchange,提问作者WashichawbachaW

