如何在Bash中正确去除多行Markdown注释?
Let's get this sorted out—your existing sed commands fall short either because they can't handle multi-line comments or they greedily delete too much content. Here are reliable solutions:
Why Your Previous Commands Failed
First sed command:
sed -e 's/<!--((.*?)\n?)+-->//g' $1
Sed works line-by-line by default, so it can't properly match comments that span multiple lines. Also, non-greedy matching (.*?) isn't supported in standard sed, making this ineffective for complex cases.Second command:
cat $1 | tr '\n' '\r' | sed -e 's/<!--.*-->//g' | tr '\r' '\n'
The.*in sed is greedy—it matches from the very first<!--to the very last-->in the entire file, wiping out all content in between (including non-comment text). That's why you saw the unwanted mass deletion.
Working Solutions
1. Perl (Cross-Platform, Most Reliable)
Perl handles multi-line input and context-aware matching seamlessly. Run this command:
perl -0777 -pe 's/<!--(?:[^-]|-[^-]|--[^>])*-->//g' your_input_file.txt
-0777: Reads the entire file as a single block instead of line-by-line, perfect for multi-line comments.- The regex
<!--(?:[^-]|-[^-]|--[^>])*-->precisely matches comments without getting tripped up by embedded<!--sequences:[^-]: Any character that's not a hyphen-[^-]: A hyphen followed by something that's not a hyphen--[^>]: Two hyphens followed by something that's not a closing>
2. GNU Sed (Linux-Specific)
If you're using GNU Sed (common on Linux systems), you can use its null-separator mode to handle multi-line content:
sed -z 's/<!--\([^-]\|-[^-]\|--[^>]\)*-->//g' your_input_file.txt
-z: Treats the input as a stream separated by null characters, effectively reading the whole file at once.- The regex works the same way as the Perl version, ensuring we only match complete comment blocks.
Test It Out
For your sample input:
some text some text
Both commands will output:
some text some text
For a multi-line comment with embedded content:
hello
inside it -->
world
You'll get:
hello world
内容的提问来源于stack exchange,提问作者Sati

