正则表达式[a-z]与[\a-z]的差异:为何后者匹配空格而前者不匹配?
[a-z] and [\a-z] Let’s break this down step by step—this is a great question that gets into how regex interprets escape sequences and character ranges inside square brackets.
[a-z] — Matching Only Lowercase Letters
[a-z] is a standard, straightforward character class: it matches any single lowercase letter from a to z. Inside square brackets, the hyphen - acts as a range operator when placed between two characters. Here, it uses the ASCII values of a (97) and z (122) to define the range. Since the space character has an ASCII value of 32—way below 97—it doesn’t fall into this range, so it never gets matched by [a-z].
[\a-z] — A Broad Range That Includes Space
This is where the confusion happens, and all of it ties to the \a escape sequence:
- In nearly all regex engines,
\ais a special escape code representing the ASCII BEL character (a control/alert character with an ASCII value of 7). - When you write
[\a-z], the hyphen again acts as a range operator—but this time, it creates a range from the character represented by\a(ASCII 7) all the way up toz(ASCII 122). - The space character (ASCII 32) sits right in the middle of 7 and 122, so it’s included in this range. Along with space, this character class will match every ASCII character from 7 to 122: that includes control characters like BEL and backspace, punctuation, numbers, uppercase letters, and lowercase letters.
A Quick Note on Common Missteps
If you actually intended to match the literal backslash + a plus lowercase letters, you’d need to escape the backslash (like [\\a-z]) or move the hyphen to the end of the character class (like [\a-z-]) to stop it from being treated as a range operator. But as written, [\a-z] uses \a as a control character, resulting in that wide range that includes space.
内容的提问来源于stack exchange,提问作者Kenny Omega

