如何修正正则表达式以匹配含多空格的目标URL(如www.d em m.)
Fixing Regex to Match URLs with Spaces (Including
www.d em m.) First, let's break down why your current regex isn't matching www.d em m.:
- HTML Entity Instead of Literal
&: Your regex uses&which matches the string&instead of the actual&character—this is a common slip-up when copying regex from HTML contexts. - Missing Lowercase Letters: The character classes only include
A-Z(uppercase), so lowercase letters liked,e,maren't being matched unless you use the case-insensitive flag. - Case Sensitivity: Without the
iflag, the regex ignores lowercase URL components entirely.
Corrected Regex
Here's the fixed version that addresses these gaps and will match your target URL www.d em m.:
(?:(?:https?|ftp|file):\/\/|www\.|ftp\.)[-a-zA-Z0-9+&@#\/%=~_|$?!:,.| ]*[a-zA-Z0-9+&@#\/%=~_|$]\.
Or, use the case-insensitive (i) flag to simplify the character classes (avoids repeating a-z):
/(?:(?:https?|ftp|file):\/\/|www\.|ftp\.)[-A-Z0-9+&@#\/%=~_|$?!:,.| ]*[A-Z0-9+&@#\/%=~_|$]\./i
Key Changes Explained
- Replaced
&with&: Now the regex correctly matches the&character in URLs. - Added Lowercase Letters: Explicitly included
a-zin character classes (or used theiflag) to cover lowercase components. - Preserved Space in Character Class: The space remains allowed in the middle section to handle URLs with spaces between characters.
Testing the Regex
When applied to www.d em m., here's how it works:
- Starts with
www.which hits the initialwww\.group. - The middle section
d em mis fully covered by the allowed characters (including spaces and lowercase letters). - Ends with
m.wheremmatches the final character class, followed by the required dot.
Optional Adjustment
If you need to allow URLs that end with a space before the dot (e.g., www.d em m .), modify the final part to permit spaces between the last valid character and the dot:
/(?:(?:https?|ftp|file):\/\/|www\.|ftp\.)[-A-Z0-9+&@#\/%=~_|$?!:,.| ]*[A-Z0-9+&@#\/%=~_|$](?:\s*)\./i
内容的提问来源于stack exchange,提问作者Alen.Toma
相关产品推荐
相关产品推荐

