使用BeautifulSoup获取下一页URL时报错:'NoneType'对象无属性'a'
错误修复与代码优化方案
一、错误原因与修复
你遇到的'NoneType' object has no attribute 'a'错误,是因为爬取到最后一页时,页面不存在<li class="next">元素,soup.find('li', {'class': 'next'})会返回None,此时调用.a自然会报错。
修复方法:先判断是否存在下一页按钮,存在才更新URL,否则直接终止循环:
next_page = soup.find('li', {'class': 'next'}) if next_page: url = url.split('/page')[0] + next_page.a['href'] # 避免多次拼接后URL格式错误 else: flag = False # 没有下一页,终止循环
二、代码优化点及完整优化后代码
你的代码存在重复DOM查询、冗余逻辑、拼写错误等问题,以下是具体优化方向和完整优化代码:
核心优化点:
- 减少重复DOM查询:原代码循环内多次调用
soup.find_all(),每次都会重新遍历DOM,优化为一次性获取所有目标节点再配对遍历。 - 简化数据库操作:用Django内置的
get_or_create()方法,替代手动判断exists()的冗余代码,一行完成「查询-不存在则创建」操作。 - 修复拼写错误:将
responce修正为正确的response。 - 优化循环逻辑:直接遍历quote和author的配对,不用按索引循环,更简洁易读。
- 增加请求异常处理:避免网络问题导致程序直接崩溃。
- 规范URL拼接:原代码直接
url += ...会导致最后一页拼接错误,改为基于基础URL拼接下一页路径。
完整优化后代码:
import requests from bs4 import BeautifulSoup from your_app.models import Quote, Author # 替换为你的实际app名称 def parser(): base_url = 'https://quotes.toscrape.com' current_url = base_url collected_count = 0 target_count = 5 # 目标抓取5条新quote while collected_count < target_count: # 处理请求异常 try: response = requests.get(current_url, timeout=10) response.raise_for_status() # 触发HTTP错误异常 except requests.exceptions.RequestException as e: print(f"请求出错: {e}") break soup = BeautifulSoup(response.text, 'html.parser') # 一次性获取所有quote和author节点,配对遍历 quotes = soup.find_all('span', class_='text') authors = soup.find_all('small', class_='author') for quote_text, author_name in zip(quotes, authors): if collected_count >= target_count: break # 检查quote是否已存在 if not Quote.objects.filter(quote=quote_text.string).exists(): # 用get_or_create简化作者的查询/创建 author, _ = Author.objects.get_or_create(name=author_name.string) Quote.objects.create(quote=quote_text.string, author=author) collected_count += 1 # 处理下一页逻辑 next_page = soup.find('li', class_='next') if next_page: current_url = base_url + next_page.a['href'] else: print("已到最后一页,终止爬取") break
补充说明
优化后的代码逻辑更清晰,执行效率更高,同时增加了容错机制,避免因网络波动或页面结构变化导致的崩溃。
内容的提问来源于stack exchange,提问作者StasFQ
相关产品推荐
相关产品推荐

