You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python正则表达式提取HTML中不含片段的URL链接

提取HTML中所有合法链接并移除片段标识符的Python实现

正则表达式设计思路

要覆盖所有合法链接类型(绝对URL、协议相对路径、根路径、相对路径),同时兼容href属性的单双引号包裹情况,最后移除#及之后的片段内容。

代码实现

import re

def extract_clean_links(html_text):
    # 匹配href属性,兼容单双引号,忽略大小写(应对HREF大写的场景)
    href_regex = re.compile(r'href=(["\'])(.*?)\1', re.IGNORECASE)
    cleaned_links = []
    
    for match in href_regex.finditer(html_text):
        raw_link = match.group(2)
        # 移除#及之后的所有内容
        clean_link = re.sub(r'#.*$', '', raw_link)
        # 过滤空链接和仅为片段的无效链接
        if clean_link.strip():
            cleaned_links.append(clean_link.strip())
    
    return cleaned_links

# 测试示例
test_html = """
<a href="https://example.com/article">HTTPS绝对链接</a>
<a href='http://example.com/post'>HTTP绝对链接</a>
<a href="//cdn.example.com/image.jpg">协议相对链接</a>
<a href="/about">根相对路径</a>
<a href="./contact">当前目录相对路径</a>
<a href="../archive">父目录相对路径</a>
<a href="readme.md">直接文件链接</a>
<a href="#footer">纯片段链接(会被过滤)</a>
<a href="/faq#section1">带片段的链接(移除片段)</a>
"""

# 提取并输出结果
for link in extract_clean_links(test_html):
    print(link)

代码说明

  1. 正则匹配href属性:href=(["\'])(.*?)\1 匹配所有用单/双引号包裹的href值,.*?采用非贪婪模式避免误匹配多个引号间的内容。
  2. 移除片段标识符:用re.sub(r'#.*$', '', raw_link)删除#及之后的所有字符,不管#在链接的哪个位置。
  3. 过滤无效链接:通过strip()和非空判断,过滤掉纯片段、空链接等无效内容。

输出结果

运行上述代码后,会输出清理后的所有合法链接:

https://example.com/article
http://example.com/post
//cdn.example.com/image.jpg
/about
./contact
../archive
readme.md
/faq

内容的提问来源于stack exchange,提问作者Mahamodul Shakil

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 19:40:56