You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬虫初学者关于find_all使用及get('href')报错的问题求助

BeautifulSoup爬虫使用问题解答

问题1:调用find_all('a')后使用.get('href')报错的原因

  • 核心原因:find_all()方法返回的是ResultSet可迭代对象,结构等同于列表,存储所有匹配到的a标签对象,只有单个标签对象才有.get()方法,列表本身不存在该方法,直接调用会触发属性报错。
  • 正确实现代码:
import requests
from bs4 import BeautifulSoup
url = 'https://example/'
page = requests.get(url)
soup = BeautifulSoup(page.text, 'lxml')
# 获取所有a标签
a_tags = soup.find_all('a')
# 遍历提取每个a标签的href属性
for a in a_tags:
    href = a.get('href')
    print(href)

问题2:多层查找的正确实现方式

  • 核心逻辑:find_all()返回的是多元素集合,需要先遍历集合拿到单个元素,再对单个元素调用find()/find_all()做二次查找,不能直接对集合本身调用查找方法。
  • 先提取h2标签再查找内部a标签的实现代码:
# 先获取所有h2标签
h2_tags = soup.find_all('h2')
# 遍历每个h2标签,查找内部的a标签
for h2 in h2_tags:
    a_tag = h2.find('a')
    # 增加判空逻辑避免h2下无a标签时报错
    if a_tag:
        href = a_tag.get('href')
        print(href)

内容的提问来源于stack exchange,提问作者Shovo Murad

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 14:57:02