You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup从两种不同结构的HTML中提取链接?

提取两种HTML结构中的目标链接

可以通过文本内容定位,找到包含"Service request"的元素,再向上或向下查找对应的<a>标签,不管嵌套关系。

示例代码

from bs4 import BeautifulSoup

# 两种测试HTML结构
html1 = '<a href="https://example.com/req1"><div>Service request</div></a>'
html2 = '<div>Service request <a href="https://example.com/req2">查看详情</a></div>'

# 通用提取函数
def get_service_request_link(html_content):
    soup = BeautifulSoup(html_content, 'html.parser')
    # 定位包含目标文本的节点
    target_text_nodes = soup.find_all(string=lambda text: text and 'Service request' in text.strip())
    for node in target_text_nodes:
        # 先找父级a标签(对应a包裹div的情况),找不到就找后续的a标签(对应div包含a的情况)
        link_tag = node.find_parent('a') or node.find_next('a')
        if link_tag and link_tag.has_attr('href'):
            return link_tag['href']
    return None

# 测试验证
print(get_service_request_link(html1))  # 输出: https://example.com/req1
print(get_service_request_link(html2))  # 输出: https://example.com/req2

逻辑说明

  • 用find_all配合lambda表达式精准定位包含目标文本的节点,避免标签嵌套差异的影响
  • 优先从目标文本节点的父级查找<a>标签,适配第一种嵌套结构
  • 若父级无<a>标签,则查找文本节点之后的第一个<a>标签,适配第二种嵌套结构
  • 最后校验<a>标签是否存在href属性,确保返回有效链接

内容的提问来源于stack exchange,提问作者Tonin thomas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 08:42:34