You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用BeautifulSoup提取指定标签之后的所有a链接元素

你需要使用BeautifulSoup提供的find_all_next()方法,该方法会提取当前节点之后所有符合条件的DOM节点,不受层级限制,完全满足你的需求。

完整实现代码

from bs4 import BeautifulSoup

html1 = """<html>
<head></head>
<body>
<p>Hello World!</p>
<a href='whatevs.com'>whatevs</a>
<p>Howdy!</p>
<a href='well.com'>well</a>
<div><span>haha</span><a href='haha.com'>haha</a></div>
<a href='goodbye.com'>Goodbye!</a>
</body>
</html>"""

# 初始化解析器
soup = BeautifulSoup(html1, 'html.parser')
# 定位目标p标签,旧版本BeautifulSoup可将string参数替换为text
howdy = soup.find('p', string='Howdy!')
# 提取所有p标签之后的a标签
target_a_list = howdy.find_all_next('a')
# 提取链接文本
text_result = [a.get_text(strip=True) for a in target_a_list]

print(target_a_list)
print(text_result)

输出结果

[<a href="well.com">well</a>, <a href="haha.com">haha</a>, <a href="goodbye.com">Goodbye!</a>]
['well', 'haha', 'Goodbye!']

方法说明

  • find_next_siblings() 仅匹配与当前节点同级、位于当前节点之后的节点,无法获取嵌套在其他标签内部的元素,因此你之前的调用结果不符合预期
  • find_next() 仅返回当前节点之后第一个匹配的元素,只能拿到单个结果
  • find_all_next() 会遍历当前节点之后的全部DOM节点,无论嵌套层级,只要符合选择器条件就会被返回,是你需要的对应find_all()的后续节点查询方法。

内容的提问来源于stack exchange,提问作者user7864386

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 23:27:03