You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用BeautifulSoup单独提取div下子元素a标签的href属性值

错误原因

你的代码存在两个核心问题:

  • 筛选div时填写的类名foo和实际页面目标div的类名speaker_social_wrap不匹配
  • 仅定位到了外层div元素,没有进一步提取内部a标签的href属性,直接打印div元素自然会输出整个标签的完整内容

解决代码

如果页面只有一个speaker_social_wrap类的div,使用如下代码即可:

# 定位目标外层div
social_wrap = soup.find('div', class_='speaker_social_wrap')
# 提取div内部所有a标签
a_tag_list = social_wrap.find_all('a')
# 遍历打印每个a标签的href属性
for a in a_tag_list:
    print(a.get('href'))

如果页面存在多个speaker_social_wrap类的div,调整为遍历所有匹配div即可:

social_wrap_list = soup.find_all('div', class_='speaker_social_wrap')
for wrap in social_wrap_list:
    a_tag_list = wrap.find_all('a')
    for a in a_tag_list:
        print(a.get('href'))

补充说明

调用元素的get('属性名')方法可以安全提取对应属性值,就算对应属性不存在也不会抛出异常,符合绝大多数爬取场景的需求。

内容的提问来源于stack exchange,提问作者Raspberry Lemon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 03:36:04