JavaScript正则匹配URL查询参数但排除HTML实体问题求助
Fixing Your URL Regex to Handle
& Edge Cases Got it, let's figure out why your current regex isn't behaving as expected and fix it up to meet your requirements.
What's Wrong With the Original Pattern
Your regex /(?:(?:https?):\/\/[-A-Z0-9+@#\/%?=~_|!:,.;]*(?:&(?!lt]))*)/gi has two critical issues:
- Missing
&in the character set: The initial character class[-A-Z0-9+@#\/%?=~_|!:,.;]doesn't include&, so it stops matching as soon as it hits the&in&. That's why it only gets up tohttps://a.com?a=b&instead of including the rest of the URL. - Incorrect lookahead check: The negative lookahead
(?!lt])is checking forlt](a closing bracket afterlt), but you need to stop when&is followed bylt;. This is why it doesn't terminate early forhttps://a.com<.
The Fixed Regex
Here's a revised pattern that addresses both issues and meets your expected behavior:
/(https?:\/\/(?:(?!<).)*)/gi
How This Works
Let's break down the key parts:
https?:\/\/: Matches the standardhttp://orhttps://URL prefix.(?:(?!<).)*: This non-capturing group acts as a "stop when we see<" guard:(?!<): A negative lookahead that checks if the next characters are not<..: Matches any single character (excluding newlines; use[\s\S]instead if you need to include line breaks).*: Repeats this check for every character until we hit<or the end of the string.
If you want a stricter version that only allows valid URL characters (instead of matching any character), use this pattern instead:
/(https?:\/\/[-A-Z0-9+@#\/%?=~_|!:,.;&]*(?:&(?!lt;)[-A-Z0-9+@#\/%?=~_|!:,.;&]*)*)/gi
This version restricts matches to common URL-safe characters and explicitly handles & segments that aren't followed by lt;.
Testing the Fix
Let's verify with your test cases:
- Input:
https://a.com?a=b&c=d- Match Result:
https://a.com?a=b&c=d✅ (full URL is captured, since there's no<to stop at)
- Match Result:
- Input:
https://a.com<- Match Result:
https://a.com✅ (stops right before<as intended)
- Match Result:
内容的提问来源于stack exchange,提问作者Veera
相关产品推荐
相关产品推荐

