You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

BeautifulSoup分页最后页面报错:AttributeError问题求助

解决爬虫分页最后一页的AttributeError问题

嘿,我来帮你搞定这个分页问题!首先咱们得弄明白为啥会报错:

问题根源

当爬到最后一页时,页面上根本没有id="next-page-link"的<li>元素,所以soup.find('li', {"id": "next-page-link"})返回的是None,你再对None调用.find('a'),自然就触发AttributeError了。你之前试了Try/Except但只爬第一页,估计是异常处理的位置没放对,直接把循环给打断了。

修复思路

咱们得把下一页链接的获取逻辑改得更健壮:

  1. 先检查有没有next-page-link这个<li>标签
  2. 要是有,再去拿里面的<a>标签和href;要是没有,就说明到最后一页了,直接终止循环
  3. 顺便给你优化下请求逻辑——你原代码里重复请求了两次同一个URL,这完全没必要,浪费资源还慢

修改后的完整代码

from bs4 import BeautifulSoup
import requests
import pandas as pd

url = "https://timetochoose.co.ao/?ct_keyword&amp;ct_ct_status&amp;ct_property_type&amp;ct_beds&amp;search-listings=true&amp;ct_country=portugal&amp;ct_state&amp;ct_city&amp;ct_price_to&amp;ct_mls&amp;lat&amp;lng"
headers = {'User-Agent': 'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:77.0) Gecko/20100101 Firefox/77.0'}
anuncios_ttc = {}
anuncios_nr = 0

while True:
    # 优化:只请求一次当前页面,避免重复调用
    response = requests.get(url, headers=headers)
    print(f"当前请求URL: {url}, 响应状态码: {response.status_code}")
    soup = BeautifulSoup(response.content, 'html.parser')
    
    anuncios = soup.find_all("div", {"class": "grid-listing-info"})
    for anuncio in anuncios:  # 把循环变量改名为anuncio,避免和列表变量重名
        titles = anuncio.find("a",{"class": "listing-link"}).text.strip()
        location = anuncio.find("p",{"class": "location muted marB0"}).text.strip()
        link = anuncio.find("a",{"class": "listing-link"}).get("href")
        
        anuncios_response = requests.get(link, headers=headers)
        anuncios_soup = BeautifulSoup(anuncios_response.text, 'html.parser')
        
        conteudo = anuncios_soup.find("div", {"id":"listing-content"}).text.strip()
        preco = anuncios_soup.find("span",{"class": "listing-price"})
        preco_imo = preco.text.strip() if preco else "N/A"
        
        quartos = anuncios_soup.find("li", {"class": "row beds"})
        nr_quartos = quartos.text.strip() if quartos else "N/A"
        
        wcs = anuncios_soup.find("li", {"class": "row baths"})
        nr_wcs = wcs.text.strip() if wcs else "N/A"
        
        tipo = anuncios_soup.find("li", {"class": "row property-type"})
        tipo_imo = tipo.text.strip() if tipo else "N/A"
        
        bairro = anuncios_soup.find("li", {"class": "row community"})
        bairro1 = bairro.text.strip() if bairro else "N/A"
        
        ref = anuncios_soup.find("li", {"class": "row propid"}).text.strip()
        
        anuncios_nr += 1
        anuncios_ttc[anuncios_nr] = [titles, location, bairro1, preco_imo, tipo_imo, nr_quartos, nr_wcs, conteudo, ref, link]
        print(f"已爬取第{anuncios_nr}条:\nTítulo: {titles}\nLocalização: {location}\nPreço: {preco_imo}\nLink: {link}\n")
    
    # 安全获取下一页链接:先判断next-page-link是否存在
    next_li = soup.find('li', {"id": "next-page-link"})
    if next_li:
        url_tag = next_li.find('a')
        if url_tag and url_tag.get('href'):
            url = url_tag.get('href')
            print(f"准备爬取下一页: {url}")
        else:
            print("没有有效下一页链接,终止循环")
            break
    else:
        print("已到达最后一页,终止循环")
        break

print(f"Nr Total de Anuncios: {anuncios_nr}")
anuncios_ttc_df = pd.DataFrame.from_dict(anuncios_ttc, orient='index', columns=['Titulo', 'Localização', 'Bairro', 'Preço', 'Tipo', 'Quartos', 'WCs', 'Descrição', 'Referência', 'Ligação'])
anuncios_ttc_df.to_csv('ttc_python.csv', index=False)  # 添加index=False避免保存额外的索引列
print("数据已成功保存到ttc_python.csv")

几个关键改进点

  • 减少重复请求:原代码里连续两次调用requests.get(url),现在只请求一次,既省时间又减轻服务器压力
  • 健壮的分页判断:先检查<li id="next-page-link">是否存在,再去获取<a>标签,彻底杜绝NoneType错误
  • 变量名优化:把循环变量从anuncios改成anuncio,避免变量名冲突导致的奇怪问题
  • 数据清洗:给所有文本字段加上.strip(),去掉多余的空格、换行,让导出的CSV更整洁
  • 调试友好:添加了更清晰的日志输出,方便你跟踪爬虫的进度和状态

这样修改后,爬虫就能正常遍历所有分页,直到最后一页自动停止,不会再报错,也不会只爬第一页啦!

内容的提问来源于stack exchange,提问作者theprodigy83

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 09:37:40