Stack Overflow首帖:如何用Grep/Sed提取跨多行的{}或[]间内容?
Hey there! Welcome to Stack Overflow—glad you're here asking questions as a Bash beginner, that's the best way to learn. Let's tackle your problem of pulling out text between curly braces {} and square brackets [], even when that content spans multiple lines.
Using Grep (Perl-Compatible Mode)
Grep usually works line-by-line, but we can tweak it to handle multi-line content with the -z flag (treats the entire input as a single "stream" using null bytes as separators) and -P (enables Perl-compatible regular expressions).
Extract content between curly braces {}
grep -zPo '(?<=\{).*?(?=\})' your_input_file.txt
-z: Processes the whole file as one continuous block instead of line-by-line-P: Lets us use lookaround assertions to target the content between delimiters(?<=\{): Positive lookbehind—locks onto the position right after a{.*?: Non-greedy match (stops at the first closing}instead of the last one, avoiding over-matching)(?=\}): Positive lookahead—locks onto the position right before a}
Extract content between square brackets []
Just swap the braces for brackets in the regex:
grep -zPo '(?<=\[).*?(?=\])' your_input_file.txt
Bonus: Keep the surrounding braces/brackets
If you want to include the delimiters in the output, skip the lookarounds and match the entire block directly:
# For braces grep -zPo '\{.*?\}' your_input_file.txt # For brackets grep -zPo '\[.*?\]' your_input_file.txt
Using Sed (Multi-Line Mode)
Sed can handle multi-line content by reading lines into its "pattern space" until it finds the closing delimiter. Here's how to set it up:
Extract content between curly braces {}
sed -n '/{/{:loop;N;/}/!b loop;s/.*{\(.*\)}.*/\1/p}' your_input_file.txt
-n: Suppresses default output so we only print the content we care about/\{/: Starts processing when a line containing{is found:loop: Creates a label to loop back to if we haven't found the closing delimiter yetN: Reads the next line into the pattern space (combines lines)/}/!b loop: If the pattern space doesn't have a}, jump back toloopand read another lines/.*{\(.*\)}.*/\1/p: Replaces everything before{and after}with just the middle content, then prints it
Extract content between square brackets []
We need to escape the brackets since they're special characters in sed:
sed -n '/\[/\{:loop;N;/\]/!b loop;s/.*\[\(.*\)\].*/\1/p}' your_input_file.txt
A Quick Heads-Up
These commands work great for non-nested braces/brackets. If you have nested structures (like { outer { inner } }), regular expressions (and thus grep/sed) will struggle because they can't handle recursive patterns. For nested cases, you'd want to use a tool like awk with custom logic or a language with proper parsing capabilities—but that's a more advanced topic for later!
Hope these solutions help you out—happy scripting!
内容的提问来源于stack exchange,提问作者sebXIII1

