grep处理含@字符串的异常行为及原因问询
Let's break down what's happening here—this is a super common regex gotcha with grep!
The core issue is that grep doesn't require your regex to match the entire input string by default. It scans for any substring within the line that fits your pattern, and the -o option tells it to output just those matching substrings instead of the whole line. You were expecting only full-string matches to show up, but without adding anchor characters to your regex, grep will happily latch onto valid partial segments of your input.
Let's walk through each of your functions to unpack the results:
funct1 (expected output)
Your regex[[:alnum:]]*@[[:alpha:]]*\.[[:alpha:]]{2,}perfectly matches the entirelocal1@domain.tldstring—local1is alphanumeric, so the first character class covers it fully, and the rest of the string lines up with the rest of the regex. No surprises here.funct2 (unexpected
@domain.tldoutput)
Your regex starts with[[:alpha:]]*, butlocal1has a digit (1) which isn't included in[[:alpha:]]. Instead of failing entirely, though, the*quantifier allows the[[:alpha:]]*part to match zero characters (since*means "0 or more"). Then the rest of the regex (@[[:alpha:]]*\.[[:alpha:]]{2,}) matches the@domain.tldsegment of your input.grep -ograbs that valid substring and outputs it.funct3 (expected no output)
Your regex[[:alpha:]]*@[[:alpha:]]*@[[:alpha:]]*\.[[:alpha:]]{2,}can't find any matching substring inlocal1@dom1@ain.tld—bothlocal1anddom1have digits, which[[:alpha:]]*can't match, and there's no way to skip those parts while still satisfying the double-@pattern. So no matches, which aligns with your expectation.funct4 (unexpected
@dom1@ain.tldoutput)
Similar to funct2, the leading[[:alpha:]]*matches zero characters (sincelocal1has a digit). The rest of the regex (@[[:alnum:]]*@[[:alpha:]]*\.[[:alpha:]]{2,}) perfectly matches the@dom1@ain.tldsegment—dom1is alphanumeric, so it fits[[:alnum:]]*, and the rest of the string checks out.grep -ooutputs this valid substring.
Fixing the behavior
If you want grep to only output results when the entire input string matches your regex, add the ^ (start-of-line anchor) and $ (end-of-line anchor) to your pattern. For example, modify funct2's regex like this:
regex="^[[:alpha:]]*@[[:alpha:]]*\.[[:alpha:]]{2,}$"
Now the regex has to match from the very start to the very end of the input line. local1@domain.tld will fail the match (because local1 has a digit that [[:alpha:]]* can't cover), so nothing gets output—exactly what you expected.
A quick reminder: the * quantifier is greedy, but crucially, it allows zero matches. That's why it can "skip" parts of your input that don't fit the character class and still find valid substrings later on.
内容的提问来源于stack exchange,提问作者SaffronSalt

