如何用grep限定符提取目标词汇?优化命令获纯单词结果
Optimized Grep Command to Extract Pure Words
Here are a couple of practical solutions to pull just the clean words from your Markdown file, stripping out the ** markers and any accompanying释义 content:
Solution 1: Using Perl-Compatible Grep (grep -P)
If your system supports grep -P (most modern Linux/macOS setups do), this single command will directly extract the words you need:
grep -oP '\*\*\K[A-Za-z]+(?=\*\*|[: ])' Principles_20_LifePrinciples.md | grep -v -e "Origin" -e "Etymology"
Breakdown:
\*\*\K: Skips the opening**without including it in the output.[A-Za-z]+: Matches one or more alphabetic characters (your target words).(?=\*\*|[: ]): A positive lookahead that stops matching when it hits closing**or a colon/space (assuming释义 starts here).- The final
grep -vfilters out unwanted entries like "Origin" and "Etymology".
Solution 2: Using Grep + Sed (Portable Across Systems)
If grep -P isn't available, this pipeline works with standard grep and sed on almost any system:
grep -o '\*\*[^*]*\*\*' Principles_20_LifePrinciples.md | grep -v -e "Origin" -e "Etymology" | sed -E 's/\*\*([A-Za-z]+).*\*\*/\1/'
Breakdown:
- The initial
grepextracts all text wrapped in**. - The second
grepremoves lines containing "Origin" or "Etymology". sed -Ecaptures the first sequence of letters inside the**and replaces the entire line with just that word, discarding any释义 content.
Both commands will give you a clean list of pure words like circumstance, case, condition, Anxiety, anxiety.
内容的提问来源于stack exchange,提问作者user9062604
相关产品推荐
相关产品推荐

