BeautifulSoup分页最后页面报错:AttributeError问题求助
解决爬虫分页最后一页的AttributeError问题
嘿,我来帮你搞定这个分页问题!首先咱们得弄明白为啥会报错:
问题根源
当爬到最后一页时,页面上根本没有id="next-page-link"的<li>元素,所以soup.find('li', {"id": "next-page-link"})返回的是None,你再对None调用.find('a'),自然就触发AttributeError了。你之前试了Try/Except但只爬第一页,估计是异常处理的位置没放对,直接把循环给打断了。
修复思路
咱们得把下一页链接的获取逻辑改得更健壮:
- 先检查有没有
next-page-link这个<li>标签 - 要是有,再去拿里面的
<a>标签和href;要是没有,就说明到最后一页了,直接终止循环 - 顺便给你优化下请求逻辑——你原代码里重复请求了两次同一个URL,这完全没必要,浪费资源还慢
修改后的完整代码
from bs4 import BeautifulSoup import requests import pandas as pd url = "https://timetochoose.co.ao/?ct_keyword&ct_ct_status&ct_property_type&ct_beds&search-listings=true&ct_country=portugal&ct_state&ct_city&ct_price_to&ct_mls&lat&lng" headers = {'User-Agent': 'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:77.0) Gecko/20100101 Firefox/77.0'} anuncios_ttc = {} anuncios_nr = 0 while True: # 优化:只请求一次当前页面,避免重复调用 response = requests.get(url, headers=headers) print(f"当前请求URL: {url}, 响应状态码: {response.status_code}") soup = BeautifulSoup(response.content, 'html.parser') anuncios = soup.find_all("div", {"class": "grid-listing-info"}) for anuncio in anuncios: # 把循环变量改名为anuncio,避免和列表变量重名 titles = anuncio.find("a",{"class": "listing-link"}).text.strip() location = anuncio.find("p",{"class": "location muted marB0"}).text.strip() link = anuncio.find("a",{"class": "listing-link"}).get("href") anuncios_response = requests.get(link, headers=headers) anuncios_soup = BeautifulSoup(anuncios_response.text, 'html.parser') conteudo = anuncios_soup.find("div", {"id":"listing-content"}).text.strip() preco = anuncios_soup.find("span",{"class": "listing-price"}) preco_imo = preco.text.strip() if preco else "N/A" quartos = anuncios_soup.find("li", {"class": "row beds"}) nr_quartos = quartos.text.strip() if quartos else "N/A" wcs = anuncios_soup.find("li", {"class": "row baths"}) nr_wcs = wcs.text.strip() if wcs else "N/A" tipo = anuncios_soup.find("li", {"class": "row property-type"}) tipo_imo = tipo.text.strip() if tipo else "N/A" bairro = anuncios_soup.find("li", {"class": "row community"}) bairro1 = bairro.text.strip() if bairro else "N/A" ref = anuncios_soup.find("li", {"class": "row propid"}).text.strip() anuncios_nr += 1 anuncios_ttc[anuncios_nr] = [titles, location, bairro1, preco_imo, tipo_imo, nr_quartos, nr_wcs, conteudo, ref, link] print(f"已爬取第{anuncios_nr}条:\nTítulo: {titles}\nLocalização: {location}\nPreço: {preco_imo}\nLink: {link}\n") # 安全获取下一页链接:先判断next-page-link是否存在 next_li = soup.find('li', {"id": "next-page-link"}) if next_li: url_tag = next_li.find('a') if url_tag and url_tag.get('href'): url = url_tag.get('href') print(f"准备爬取下一页: {url}") else: print("没有有效下一页链接,终止循环") break else: print("已到达最后一页,终止循环") break print(f"Nr Total de Anuncios: {anuncios_nr}") anuncios_ttc_df = pd.DataFrame.from_dict(anuncios_ttc, orient='index', columns=['Titulo', 'Localização', 'Bairro', 'Preço', 'Tipo', 'Quartos', 'WCs', 'Descrição', 'Referência', 'Ligação']) anuncios_ttc_df.to_csv('ttc_python.csv', index=False) # 添加index=False避免保存额外的索引列 print("数据已成功保存到ttc_python.csv")
几个关键改进点
- 减少重复请求:原代码里连续两次调用
requests.get(url),现在只请求一次,既省时间又减轻服务器压力 - 健壮的分页判断:先检查
<li id="next-page-link">是否存在,再去获取<a>标签,彻底杜绝NoneType错误 - 变量名优化:把循环变量从
anuncios改成anuncio,避免变量名冲突导致的奇怪问题 - 数据清洗:给所有文本字段加上
.strip(),去掉多余的空格、换行,让导出的CSV更整洁 - 调试友好:添加了更清晰的日志输出,方便你跟踪爬虫的进度和状态
这样修改后,爬虫就能正常遍历所有分页,直到最后一页自动停止,不会再报错,也不会只爬第一页啦!
内容的提问来源于stack exchange,提问作者theprodigy83
相关产品推荐
相关产品推荐

