基于BeautifulSoup的租房/购房爬虫爬取多页后异常求助
解决爬虫爬取页面重复、数据量下降及分批写入Excel的问题
首先澄清一下:列表过载不是导致这个问题的核心原因——Python完全能轻松处理几万条数据的列表,不会因此引发页面重复加载或爬取数据锐减的情况。你的问题大概率是网站反爬机制在限制请求,下面我会先解决这个核心问题,再帮你实现分批写入Excel的方案。
一、先搞定反爬导致的重复页面/数据减少问题
网站检测到频繁的自动化请求后,会返回重复页面内容(甚至空页面),这就是你看到strona=34重复输出、每页仅爬1条数据的原因。可以通过以下方式缓解:
- 添加随机请求延迟,模拟人类浏览节奏
- 使用随机User-Agent,避免固定标识被识别为爬虫
- 检查响应状态码,确保请求成功,失败时尝试重试
二、实现分批写入Excel(解决你担心的内存占用问题)
如果担心大量数据占用内存,可以每爬取一定页数就把当前数据写入Excel,清空列表后继续爬取,后续写入用追加模式即可。
修改后的完整代码
from bs4 import BeautifulSoup from requests import get import pandas as pd import time import random from fake_useragent import UserAgent # 需先安装:pip install fake-useragent # 初始化随机UA生成器 ua = UserAgent() # 定义每次写入的批次大小(比如每20页写一次) BATCH_SIZE = 20 # 初始化临时存储单批次数据的列表 temp_data = [] # 估算总页数:9500条广告,按每页20条计算约475页 for page in range(1, 476): # 生成随机User-Agent headers = {'User-Agent': ua.random} link = f'https://ogloszenia.trojmiasto.pl/nieruchomosci-rynek-pierwotny/?strona={page}' try: r = get(link, headers=headers) # 检查请求是否成功 if r.status_code != 200: print(f"请求页面{page}失败,状态码:{r.status_code},重试中...") time.sleep(5) r = get(link, headers=headers) if r.status_code != 200: print(f"页面{page}重试仍失败,跳过") continue zupa = BeautifulSoup(r.text, 'html.parser') ogloszenia = zupa.find_all('div', class_="list__item") print(f"正在处理页面:{link},共找到{len(ogloszenia)}条广告") for ogl in ogloszenia: # 简化数据提取逻辑,避免多层try-except tytul = ogl.find('h2', class_="list__item__content__title").text.strip() if ogl.find('h2', class_="list__item__content__title") else '' details = ogl.find_all('p', class_="list__item__details__icons__element__desc") powierzchnia = details[0].text.strip() if len(details)>=1 else '' liczba_pok = details[1].text.strip() if len(details)>=2 else '' pietro = details[2].text.strip() if len(details)>=3 else '' rok = details[3].text.strip() if len(details)>=4 else '' lokalizacja = ogl.find('p', class_="list__item__content__subtitle").text.strip() if ogl.find('p', class_="list__item__content__subtitle") else '' cena = ogl.find('p', class_="list__item__price__value").text.strip() if ogl.find('p', class_="list__item__price__value") else '' cena_m = ogl.find('p', class_="list__item__details__info details--info--price").text.strip() if ogl.find('p', class_="list__item__details__info details--info--price") else '' # 将单条数据存入临时列表 temp_data.append({ "Tytul": tytul, "Lokalizacja": lokalizacja, "Cena": cena, "Cena za metr": cena_m, "Powierzchnia": powierzchnia, "Pietro": pietro, "Liczba pokoi": liczba_pok, "Rok budowy": rok }) # 达到批次大小,执行写入操作 if page % BATCH_SIZE == 0: df = pd.DataFrame(temp_data) # 第一次写入创建文件,后续写入追加内容 if page == BATCH_SIZE: df.to_excel('rynek_pierwotnyy.xls', index=False) else: with pd.ExcelWriter('rynek_pierwotnyy.xls', mode='a', if_sheet_exists='overlay') as writer: df.to_excel(writer, index=False, header=False) print(f"已写入第{page-BATCH_SIZE+1}至{page}页数据到Excel") # 清空临时列表,准备存储下一批数据 temp_data = [] # 随机延迟1-3秒,避免触发反爬 time.sleep(random.uniform(1, 3)) except Exception as e: print(f"处理页面{page}时出错:{str(e)}") time.sleep(5) # 处理最后一批不足BATCH_SIZE的数据 if temp_data: df = pd.DataFrame(temp_data) with pd.ExcelWriter('rynek_pierwotnyy.xls', mode='a', if_sheet_exists='overlay') as writer: df.to_excel(writer, index=False, header=False) print("已写入最后一批数据到Excel")
关键修改点说明
- 反爬优化:
- 使用
fake-useragent生成随机UA,避免固定标识被识别 - 添加随机延迟,模拟人类浏览间隔
- 增加状态码检查,失败时重试,降低丢数据概率
- 使用
- 分批写入:
- 用
temp_data临时存储单批次数据,达到批次大小后写入Excel - 第一次写入创建文件,后续用
mode='a'追加内容,避免覆盖之前的数据 - 单独处理最后一批剩余数据,确保数据完整
- 用
- 代码简化:
- 用
find替代find_all[0],减少冗余代码和IndexError风险 - 用f-string拼接URL,可读性更高
- 移除无意义的
sys.getsizeof(tytuly)代码
- 用
内容的提问来源于stack exchange,提问作者aadams
相关产品推荐
相关产品推荐

