如何在Bash中用AWK匹配并打印所有符合规则的句子
Solution to Extract All Valid Sentences with Awk
Got it, let's fix this! The issue with your current awk command is that match() only finds the first matching sentence in the input string—without a loop to keep searching the remaining text, you'll only ever get that first result. Here's how to modify it to capture all valid sentences that meet your criteria:
Version 1: Output all valid sentences in a single line (matches your expected output)
text="incorrect sentence! this is not sentence! This is sentence. this is not sencence. This is correct sentence." echo "$text" | awk '{ output = "" while (match($0, /[A-Z][^!?.]*[!?.]/)) { # Extract the matched valid sentence sentence = substr($0, RSTART, RLENGTH) # Append to output with space separation output = output (output ? " " : "") sentence # Trim the input string to start after the current match $0 = substr($0, RSTART + RLENGTH) } print output }'
Output for your second example:
This is sentence. This is correct sentence.
Version 2: Output each valid sentence on its own line
If you prefer individual lines instead of a single concatenated string, use this simpler loop:
echo "$text" | awk '{ while (match($0, /[A-Z][^!?.]*[!?.]/)) { print substr($0, RSTART, RLENGTH) $0 = substr($0, RSTART + RLENGTH) } }'
Output for your first example:
This is sentence!
This is sentence.
How it works:
- The
while (match(...))loop runs repeatedly as long as a valid sentence is found in the remaining text. substr($0, RSTART, RLENGTH)uses the positions stored bymatch()to extract the exact matched sentence.$0 = substr($0, RSTART + RLENGTH)updates the input string to skip the part we've already processed, so the next iteration looks for the next valid sentence.- The regex
/[A-Z][^!?.]*[!?.]/specifically targets strings that:- Start with an uppercase letter
- Follow with any characters except
!,?, or. - End with one of those three punctuation marks
内容的提问来源于stack exchange,提问作者Amateur
相关产品推荐
相关产品推荐

