基于同一输入文件的嵌套grep模式搜索与多输出文件生成问询
Hey there! Let's tackle this problem—since you need to first locate main chapters (E:Chapter) and then check for subchapters (Sub Chapter) under each one to extract specific content, plain grep alone won't cut it (it can't handle the "contextual dependency" between main and subchapters). Here are two solid solutions tailored to your needs:
Awk is perfect for this kind of line-by-line contextual processing. This script will automatically track each main chapter, check for subchapters under it, and write the subchapter content (plus 8 following lines) to a dedicated file like chapter1.txt:
BEGIN { current_chapter = "" in_subchapter = 0 sub_lines_remaining = 0 IGNORECASE = 1 # Uncomment this if you want case-insensitive matches (e.g., "Sub chapter" works too) } # Match lines with main chapters (E:ChapterX) /.*E:chapter.*/ { # Extract the chapter number/name for the output file current_chapter = gensub(/.*E:chapter([0-9]+).*/, "\\1", "g", $0) # Reset subchapter tracking for the new main chapter in_subchapter = 0 sub_lines_remaining = 0 next } # Match subchapters only if we're currently tracking a main chapter current_chapter != "" && /.*Sub[[:space:]]Chapter.*/ { in_subchapter = 1 sub_lines_remaining = 8 # Keep 8 lines after the subchapter # Write the subchapter line to the corresponding file print $0 > "chapter" current_chapter ".txt" next } # Write subsequent lines if we're in subchapter extraction mode in_subchapter == 1 && sub_lines_remaining > 0 { print $0 > "chapter" current_chapter ".txt" sub_lines_remaining-- # Exit subchapter mode once we've written all 8 lines if (sub_lines_remaining == 0) { in_subchapter = 0 } }
How to use it:
Run this command in your terminal:
awk -f chapter_extractor.awk input.txt
- It will skip any main chapters that don't have a subchapter under them.
- If you don't want case-insensitive matching, just remove the
IGNORECASE = 1line.
If you prefer sticking to basic shell tools, this approach uses grep to find main chapter positions, then checks each one for subchapters:
Step 1: Extract main chapter line numbers
First, get the line numbers of all main chapters:
grep -n "E:chapter" input.txt | cut -d: -f1 > chapter_lines.txt
Step 2: Loop through each main chapter and check for subchapters
while read -r start_line; do # Get the chapter name (e.g., "chapter1" from "E:chapter1") chapter_name=$(sed -n "${start_line}p" input.txt | grep -o "E:chapter.*" | sed 's/E://') # Check if there's a subchapter within the next 20 lines of the main chapter (adjust 20 as needed) has_sub=$(sed -n "${start_line},+20p" input.txt | grep -q "Sub[[:space:]]Chapter" && echo "yes" || echo "no") if [ "$has_sub" = "yes" ]; then # Find the start line of the next main chapter (or end of file if it's the last one) next_start=$(grep -n "E:chapter" input.txt | awk -v current="$start_line" '$1 > current {print $1; exit}') [ -z "$next_start" ] && next_start=$(wc -l < input.txt) # Extract the subchapter + 8 lines within the current main chapter's range sed -n "${start_line},${next_start}p" input.txt | grep -A8 "Sub[[:space:]]Chapter" > "${chapter_name}.txt" fi done < chapter_lines.txt
Quick note on your original grep command:
Your initial grep -rwA 2 [E:Chapter] input.txt has a syntax issue: [E:Chapter] is a character class that matches any single character from E, :, C, h, a, p, t, e, r. To match the full string E:Chapter, remove the square brackets:
grep -rwA2 "E:Chapter" input.txt > output.txt
内容的提问来源于stack exchange,提问作者pixie

