You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于同一输入文件的嵌套grep模式搜索与多输出文件生成问询

Hey there! Let's tackle this problem—since you need to first locate main chapters (E:Chapter) and then check for subchapters (Sub Chapter) under each one to extract specific content, plain grep alone won't cut it (it can't handle the "contextual dependency" between main and subchapters). Here are two solid solutions tailored to your needs:

方案一:用Awk一站式处理(最灵活推荐)

Awk is perfect for this kind of line-by-line contextual processing. This script will automatically track each main chapter, check for subchapters under it, and write the subchapter content (plus 8 following lines) to a dedicated file like chapter1.txt:

BEGIN {
    current_chapter = ""
    in_subchapter = 0
    sub_lines_remaining = 0
    IGNORECASE = 1  # Uncomment this if you want case-insensitive matches (e.g., "Sub chapter" works too)
}

# Match lines with main chapters (E:ChapterX)
/.*E:chapter.*/ {
    # Extract the chapter number/name for the output file
    current_chapter = gensub(/.*E:chapter([0-9]+).*/, "\\1", "g", $0)
    # Reset subchapter tracking for the new main chapter
    in_subchapter = 0
    sub_lines_remaining = 0
    next
}

# Match subchapters only if we're currently tracking a main chapter
current_chapter != "" && /.*Sub[[:space:]]Chapter.*/ {
    in_subchapter = 1
    sub_lines_remaining = 8  # Keep 8 lines after the subchapter
    # Write the subchapter line to the corresponding file
    print $0 > "chapter" current_chapter ".txt"
    next
}

# Write subsequent lines if we're in subchapter extraction mode
in_subchapter == 1 && sub_lines_remaining > 0 {
    print $0 > "chapter" current_chapter ".txt"
    sub_lines_remaining--
    # Exit subchapter mode once we've written all 8 lines
    if (sub_lines_remaining == 0) {
        in_subchapter = 0
    }
}

How to use it:

Run this command in your terminal:

awk -f chapter_extractor.awk input.txt
  • It will skip any main chapters that don't have a subchapter under them.
  • If you don't want case-insensitive matching, just remove the IGNORECASE = 1 line.
方案二:用Grep + Shell循环(For shell-savvy users)

If you prefer sticking to basic shell tools, this approach uses grep to find main chapter positions, then checks each one for subchapters:

Step 1: Extract main chapter line numbers

First, get the line numbers of all main chapters:

grep -n "E:chapter" input.txt | cut -d: -f1 > chapter_lines.txt

Step 2: Loop through each main chapter and check for subchapters

while read -r start_line; do
    # Get the chapter name (e.g., "chapter1" from "E:chapter1")
    chapter_name=$(sed -n "${start_line}p" input.txt | grep -o "E:chapter.*" | sed 's/E://')
    # Check if there's a subchapter within the next 20 lines of the main chapter (adjust 20 as needed)
    has_sub=$(sed -n "${start_line},+20p" input.txt | grep -q "Sub[[:space:]]Chapter" && echo "yes" || echo "no")
    
    if [ "$has_sub" = "yes" ]; then
        # Find the start line of the next main chapter (or end of file if it's the last one)
        next_start=$(grep -n "E:chapter" input.txt | awk -v current="$start_line" '$1 > current {print $1; exit}')
        [ -z "$next_start" ] && next_start=$(wc -l < input.txt)
        
        # Extract the subchapter + 8 lines within the current main chapter's range
        sed -n "${start_line},${next_start}p" input.txt | grep -A8 "Sub[[:space:]]Chapter" > "${chapter_name}.txt"
    fi
done < chapter_lines.txt

Quick note on your original grep command:

Your initial grep -rwA 2 [E:Chapter] input.txt has a syntax issue: [E:Chapter] is a character class that matches any single character from E, :, C, h, a, p, t, e, r. To match the full string E:Chapter, remove the square brackets:

grep -rwA2 "E:Chapter" input.txt > output.txt

内容的提问来源于stack exchange,提问作者pixie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 07:55:15