You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python正则匹配含非HTTP(S)链接的完整<a>标签

Got it, let's fix that regex issue for you! The key here is to craft a regex that targets the entire <a> tag, specifically checking that its href value doesn't start with http:// or https://—and avoiding those accidental matches of HTTP links.

Optimized Regex Pattern

<a\s+(?:[^>]*?\s+)?href=["'](?!https?:\/\/)[^"']+["'][^>]*>.*?<\/a>

Breakdown of the Pattern

Let's walk through each part to understand how it works:

  • <a\s+: Matches the start of the anchor tag, ensuring there's at least one space after <a (to avoid false matches with similar-looking text).
  • (?:[^>]*?\s+)?: A non-capturing group that handles any extra attributes (like class, id) that might come before href. The [^>]*? is non-greedy to stop before hitting the href attribute, and the ? makes this section optional (for when href is the first attribute).
  • href=["']: Matches the start of the href attribute, supporting both single and double quotes.
  • (?!https?:\/\/): Negative lookahead—this is the critical part. It ensures the text immediately following the quote is NOT http:// or https://.
  • [^"']+: Matches the actual link value, stopping when it hits the closing quote (works for both single and double quotes).
  • ["']: Closes the href attribute quote.
  • [^>]*>: Matches any remaining attributes in the opening tag, then the closing >.
  • .*?<\/a>: Non-greedily matches all content inside the tag (text, other elements) until it hits the closing </a>—prevents accidentally matching multiple tags at once.

Test Cases

Matches (These will be captured):

<a href="example.com/index.html"> bla</a>
<a class="nav-link" href='/about'>About Us</a>
<a href="ftp://fileserver.com/docs">Download Docs</a>

Non-Matches (These will be ignored):

<a href="https://www.google.com/">bla2 </a>
<a href="http://example.org">Official Site</a>
<a>No href attribute here</a>

Notes

  • This regex handles both single and double quotes for href values.
  • The non-greedy quantifiers (*?) prevent "over-matching" (e.g., grabbing multiple <a> tags in one go).
  • It won't match <a> tags without a href attribute (since your goal is to target tags with non-HTTP links).
  • If you need to handle escaped quotes inside href values (rare in valid HTML), you'd need a minor adjustment, but this works for most real-world scenarios.

内容的提问来源于stack exchange,提问作者Akash Sundaresh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 08:59:01