正则表达式:对\b边界行为的困惑及十六进制匹配问题
Got it, let's fix this regex issue for your Atom custom syntax highlighting. The core problem here is that \b (word boundary) behaves differently for $ vs 0x—since $ isn't a word character (word chars are a-z, A-Z, 0-9, _), \b doesn't work as expected to block unwanted preceding characters for the $ prefix.
Here's the single regex that will handle both cases correctly, avoiding false matches like t0x0t:
(?<!\w)(?:\$|0x)[A-Fa-f0-9]+
Breakdown of how this works:
(?<!\w): A negative lookbehind assertion that ensures there is no word character immediately before our hex prefix. This blocks cases liket0x0t(wheretis a word character before0x) and also prevents unwanted matches likea$123(whereais a word character before$).(?:\$|0x): A non-capturing group that matches either the$or0xprefix. Using?:avoids creating an unnecessary capture group, which is cleaner for syntax highlighting.[A-Fa-f0-9]+: Matches one or more hexadecimal digits (upper/lowercase letters + numbers).
Test cases this handles correctly:
- ✅ Matches
$affe(whether at the start of a line or after a non-word character like!$affe) - ✅ Matches
0xeffa - ❌ Does NOT match
t0x0t(thetbefore0xtriggers the negative lookbehind, so the regex skips this) - ❌ Does NOT match
foo$123(theobefore$is a word character, so the lookbehind blocks the match)
Why this is better than using \b:\b relies on the transition between word and non-word characters. For 0x, \b0x works because 0 is a word character—if there's a word character before it (like t), there's no word boundary between t and 0, so it doesn't match. But for $, \b$ would actually match cases like t$affe (since t is a word character and $ is non-word, creating a boundary), which we don't want. The negative lookbehind fixes this by explicitly checking for the absence of a preceding word character, regardless of the prefix type.
内容的提问来源于stack exchange,提问作者koalag

