如何用Python正则匹配含非HTTP(S)链接的完整<a>标签
Got it, let's fix that regex issue for you! The key here is to craft a regex that targets the entire <a> tag, specifically checking that its href value doesn't start with http:// or https://—and avoiding those accidental matches of HTTP links.
Optimized Regex Pattern
<a\s+(?:[^>]*?\s+)?href=["'](?!https?:\/\/)[^"']+["'][^>]*>.*?<\/a>
Breakdown of the Pattern
Let's walk through each part to understand how it works:
<a\s+: Matches the start of the anchor tag, ensuring there's at least one space after<a(to avoid false matches with similar-looking text).(?:[^>]*?\s+)?: A non-capturing group that handles any extra attributes (likeclass,id) that might come beforehref. The[^>]*?is non-greedy to stop before hitting thehrefattribute, and the?makes this section optional (for whenhrefis the first attribute).href=["']: Matches the start of thehrefattribute, supporting both single and double quotes.(?!https?:\/\/): Negative lookahead—this is the critical part. It ensures the text immediately following the quote is NOThttp://orhttps://.[^"']+: Matches the actual link value, stopping when it hits the closing quote (works for both single and double quotes).["']: Closes thehrefattribute quote.[^>]*>: Matches any remaining attributes in the opening tag, then the closing>..*?<\/a>: Non-greedily matches all content inside the tag (text, other elements) until it hits the closing</a>—prevents accidentally matching multiple tags at once.
Test Cases
Matches (These will be captured):
<a href="example.com/index.html"> bla</a> <a class="nav-link" href='/about'>About Us</a> <a href="ftp://fileserver.com/docs">Download Docs</a>
Non-Matches (These will be ignored):
<a href="https://www.google.com/">bla2 </a> <a href="http://example.org">Official Site</a> <a>No href attribute here</a>
Notes
- This regex handles both single and double quotes for
hrefvalues. - The non-greedy quantifiers (
*?) prevent "over-matching" (e.g., grabbing multiple<a>tags in one go). - It won't match
<a>tags without ahrefattribute (since your goal is to target tags with non-HTTP links). - If you need to handle escaped quotes inside
hrefvalues (rare in valid HTML), you'd need a minor adjustment, but this works for most real-world scenarios.
内容的提问来源于stack exchange,提问作者Akash Sundaresh
相关产品推荐
相关产品推荐

