You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于BeautifulSoup的租房/购房爬虫爬取多页后异常求助

解决爬虫爬取页面重复、数据量下降及分批写入Excel的问题

首先澄清一下:列表过载不是导致这个问题的核心原因——Python完全能轻松处理几万条数据的列表,不会因此引发页面重复加载或爬取数据锐减的情况。你的问题大概率是网站反爬机制在限制请求,下面我会先解决这个核心问题,再帮你实现分批写入Excel的方案。

一、先搞定反爬导致的重复页面/数据减少问题

网站检测到频繁的自动化请求后,会返回重复页面内容(甚至空页面),这就是你看到strona=34重复输出、每页仅爬1条数据的原因。可以通过以下方式缓解:

  • 添加随机请求延迟,模拟人类浏览节奏
  • 使用随机User-Agent,避免固定标识被识别为爬虫
  • 检查响应状态码,确保请求成功,失败时尝试重试

二、实现分批写入Excel(解决你担心的内存占用问题)

如果担心大量数据占用内存,可以每爬取一定页数就把当前数据写入Excel,清空列表后继续爬取,后续写入用追加模式即可。

修改后的完整代码

from bs4 import BeautifulSoup
from requests import get
import pandas as pd
import time
import random
from fake_useragent import UserAgent  # 需先安装:pip install fake-useragent

# 初始化随机UA生成器
ua = UserAgent()

# 定义每次写入的批次大小(比如每20页写一次)
BATCH_SIZE = 20
# 初始化临时存储单批次数据的列表
temp_data = []

# 估算总页数:9500条广告,按每页20条计算约475页
for page in range(1, 476):
    # 生成随机User-Agent
    headers = {'User-Agent': ua.random}
    link = f'https://ogloszenia.trojmiasto.pl/nieruchomosci-rynek-pierwotny/?strona={page}'
    
    try:
        r = get(link, headers=headers)
        # 检查请求是否成功
        if r.status_code != 200:
            print(f"请求页面{page}失败,状态码:{r.status_code},重试中...")
            time.sleep(5)
            r = get(link, headers=headers)
            if r.status_code != 200:
                print(f"页面{page}重试仍失败,跳过")
                continue
        
        zupa = BeautifulSoup(r.text, 'html.parser')
        ogloszenia = zupa.find_all('div', class_="list__item")
        print(f"正在处理页面:{link},共找到{len(ogloszenia)}条广告")

        for ogl in ogloszenia:
            # 简化数据提取逻辑,避免多层try-except
            tytul = ogl.find('h2', class_="list__item__content__title").text.strip() if ogl.find('h2', class_="list__item__content__title") else ''
            details = ogl.find_all('p', class_="list__item__details__icons__element__desc")
            powierzchnia = details[0].text.strip() if len(details)>=1 else ''
            liczba_pok = details[1].text.strip() if len(details)>=2 else ''
            pietro = details[2].text.strip() if len(details)>=3 else ''
            rok = details[3].text.strip() if len(details)>=4 else ''
            lokalizacja = ogl.find('p', class_="list__item__content__subtitle").text.strip() if ogl.find('p', class_="list__item__content__subtitle") else ''
            cena = ogl.find('p', class_="list__item__price__value").text.strip() if ogl.find('p', class_="list__item__price__value") else ''
            cena_m = ogl.find('p', class_="list__item__details__info details--info--price").text.strip() if ogl.find('p', class_="list__item__details__info details--info--price") else ''
            
            # 将单条数据存入临时列表
            temp_data.append({
                "Tytul": tytul,
                "Lokalizacja": lokalizacja,
                "Cena": cena,
                "Cena za metr": cena_m,
                "Powierzchnia": powierzchnia,
                "Pietro": pietro,
                "Liczba pokoi": liczba_pok,
                "Rok budowy": rok
            })
        
        # 达到批次大小,执行写入操作
        if page % BATCH_SIZE == 0:
            df = pd.DataFrame(temp_data)
            # 第一次写入创建文件,后续写入追加内容
            if page == BATCH_SIZE:
                df.to_excel('rynek_pierwotnyy.xls', index=False)
            else:
                with pd.ExcelWriter('rynek_pierwotnyy.xls', mode='a', if_sheet_exists='overlay') as writer:
                    df.to_excel(writer, index=False, header=False)
            print(f"已写入第{page-BATCH_SIZE+1}至{page}页数据到Excel")
            # 清空临时列表,准备存储下一批数据
            temp_data = []
        
        # 随机延迟1-3秒,避免触发反爬
        time.sleep(random.uniform(1, 3))
    
    except Exception as e:
        print(f"处理页面{page}时出错:{str(e)}")
        time.sleep(5)

# 处理最后一批不足BATCH_SIZE的数据
if temp_data:
    df = pd.DataFrame(temp_data)
    with pd.ExcelWriter('rynek_pierwotnyy.xls', mode='a', if_sheet_exists='overlay') as writer:
        df.to_excel(writer, index=False, header=False)
    print("已写入最后一批数据到Excel")

关键修改点说明

  1. 反爬优化:
    • 使用fake-useragent生成随机UA,避免固定标识被识别
    • 添加随机延迟,模拟人类浏览间隔
    • 增加状态码检查,失败时重试,降低丢数据概率
  2. 分批写入:
    • 用temp_data临时存储单批次数据,达到批次大小后写入Excel
    • 第一次写入创建文件,后续用mode='a'追加内容,避免覆盖之前的数据
    • 单独处理最后一批剩余数据,确保数据完整
  3. 代码简化:
    • 用find替代find_all[0],减少冗余代码和IndexError风险
    • 用f-string拼接URL,可读性更高
    • 移除无意义的sys.getsizeof(tytuly)代码

内容的提问来源于stack exchange,提问作者aadams

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 05:25:12