Python爬虫初学者关于find_all使用及get('href')报错的问题求助
BeautifulSoup爬虫使用问题解答
问题1:调用find_all('a')后使用.get('href')报错的原因
- 核心原因:
find_all()方法返回的是ResultSet可迭代对象,结构等同于列表,存储所有匹配到的a标签对象,只有单个标签对象才有.get()方法,列表本身不存在该方法,直接调用会触发属性报错。 - 正确实现代码:
import requests from bs4 import BeautifulSoup url = 'https://example/' page = requests.get(url) soup = BeautifulSoup(page.text, 'lxml') # 获取所有a标签 a_tags = soup.find_all('a') # 遍历提取每个a标签的href属性 for a in a_tags: href = a.get('href') print(href)
问题2:多层查找的正确实现方式
- 核心逻辑:
find_all()返回的是多元素集合,需要先遍历集合拿到单个元素,再对单个元素调用find()/find_all()做二次查找,不能直接对集合本身调用查找方法。 - 先提取h2标签再查找内部a标签的实现代码:
# 先获取所有h2标签 h2_tags = soup.find_all('h2') # 遍历每个h2标签,查找内部的a标签 for h2 in h2_tags: a_tag = h2.find('a') # 增加判空逻辑避免h2下无a标签时报错 if a_tag: href = a_tag.get('href') print(href)
内容的提问来源于stack exchange,提问作者Shovo Murad
相关产品推荐
相关产品推荐

