You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过span文本‘View Sport’获取对应a标签的href属性值

解决方案:通过Span文本获取所属A标签的Href值

你可以通过以下两种方式实现需求:

方法一:先找Span再向上定位父级A标签

基于你现有的代码,找到匹配的span元素后,利用BeautifulSoup的父节点查找方法定位到对应的a标签,再提取href属性:

from bs4 import BeautifulSoup 
import re

text = """<td align="center" width="225" height="50" bgcolor="#2175bc" style="height:50px;display:block;font-family:'helvetica neue','helvetica','arial',sans-serif;font-size:17px;border:solid 1px #2175bc;border-radius:6px;">
                    <a href="https://trg/portal" target="_blank" style="color:#ffffff;text-decoration:none;line-height:50px;width:100%;display:inline-block;">
                    <span style="color:#ffffff;line-height:50px;">View Sport<span></span></span></a>
                </td>"""

soup = BeautifulSoup(text, 'html.parser')
              
# 匹配包含目标文本的span
target_span = soup.find("span", text=re.compile("View Sport|view this Sport", re.IGNORECASE))

if target_span:
    # 向上查找最近的a标签父节点
    parent_link = target_span.find_parent('a')
    if parent_link:
        # 提取href属性值
        print(parent_link.get('href'))  # 输出结果:https://trg/portal
    else:
        print("未找到对应的a标签")
else:
    print("未匹配到目标span元素")

关键说明

  • find_parent('a'):相比直接使用target_span.parent,这个方法更健壮,能自动跳过span和a之间可能存在的其他嵌套元素,精准定位到父级a标签。
  • get('href'):安全获取属性值,即使a标签没有href属性也不会抛出异常,会返回None。

方法二:用CSS选择器直接定位(简化写法)

如果你的BeautifulSoup版本在4.7.0及以上,可以使用内置的CSS选择器语法,一步到位找到包含目标文本span的a标签:

from bs4 import BeautifulSoup 
import re

text = """<td align="center" width="225" height="50" bgcolor="#2175bc" style="height:50px;display:block;font-family:'helvetica neue','helvetica','arial',sans-serif;font-size:17px;border:solid 1px #2175bc;border-radius:6px;">
                    <a href="https://trg/portal" target="_blank" style="color:#ffffff;text-decoration:none;line-height:50px;width:100%;display:inline-block;">
                    <span style="color:#ffffff;line-height:50px;">View Sport<span></span></span></a>
                </td>"""

soup = BeautifulSoup(text, 'html.parser')

# 直接匹配包含指定文本span的a标签
target_link = soup.select_one('a:has(span:-soup-contains-re("View Sport|view this Sport"))')
if target_link:
    print(target_link.get('href'))  # 输出结果:https://trg/portal

关键说明

  • :has():CSS选择器的逻辑,用来筛选包含指定子元素的父元素。
  • -soup-contains-re():BeautifulSoup扩展的正则匹配文本选择器,支持大小写不敏感的多文本匹配。

内容的提问来源于stack exchange,提问作者Ajay Kk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 17:05:22