网页URL正则表达式错误排查请求(邮箱/手机号正则已可用)
Let's break down the problems in your current code and fix them step by step:
1. Syntax Error in the Search Call
First off, you're passing https://.facebook.com directly to search() without wrapping it in quotes—Python will throw a syntax error here because it doesn't recognize that as a string. Also, the leading dot before facebook.com is probably a typo; valid URLs don't start with a dot after ://.
Fix this line to:
mo2 = UrlRegex.search("https://facebook.com") # Or with www prefix: mo2 = UrlRegex.search("https://www.facebook.com")
2. Flaws in the Regex Pattern
Your current regex has a few critical issues that prevent it from matching valid URLs correctly:
- The
(https://.)part uses an unescaped.which matches any single character (not the start of a domain). So it would only match something likehttps://xfacebook.com, which isn't what you want. - The pattern is too restrictive: domains can include hyphens (
-) and multiple subdomains (likewww), but your regex only allows[a-zA-Z0-9]+before the final dot. - The fragmented structure with unnecessary spaces (even with
re.VERBOSE) makes it hard to capture full valid domains.
Corrected Regex
Here's a more robust pattern that handles standard HTTPS URLs (including subdomains and common top-level domains like .com, .org, .co.uk):
import re # Updated regex to match valid HTTPS base URLs UrlRegex = re.compile( r''' (https:// # Match the required HTTPS protocol prefix [a-zA-Z0-9.-]+ # Match subdomains and main domain (allow hyphens and dots) \.[a-zA-Z]{2,}) # Match the top-level domain (2+ letters, e.g., .com, .org) ''', re.VERBOSE )
How This Works:
https://: Explicitly matches the secure protocol prefix you're targeting.[a-zA-Z0-9.-]+: Captures any combination of letters, numbers, hyphens, and dots—this covers subdomains (likewww.) and the main domain name.\.[a-zA-Z]{2,}: Matches the top-level domain, requiring at least 2 letters after the final dot to avoid invalid single-character TLDs.
3. Testing the Fixed Code
Putting it all together, your code should now work as expected:
import re UrlRegex = re.compile( r''' (https:// [a-zA-Z0-9.-]+ \.[a-zA-Z]{2,}) ''', re.VERBOSE ) mo2 = UrlRegex.search("https://www.facebook.com") if mo2: print(mo2.group()) # Output: https://www.facebook.com
Note: If you need to match more complex URLs (like those with paths, query parameters, or HTTP protocol), you can expand the regex further—this version focuses on capturing the base domain part of HTTPS URLs, which seems to be your goal.
内容的提问来源于stack exchange,提问作者leon mpaka

