优化Python房产爬虫脚本:解决封禁、超时及302/CORS报错
爬虫优化请求:解决302/CORS问题并完成全量爬取
我写了一个基础Python脚本,用来爬取房产网站数据,把房源地址和价格存入CSV文件。目标房源有5000多条,但现在脚本爬了大概2000条就超时了,控制台还出现302和CORS policy错误。我已经加了sleep(randint(1, 5))设置随机请求间隔,但还需要更多优化办法,希望能高效完成全量爬取,同时尊重目标网站、降低资源负载。(刚接触Python和爬虫,有基础错误请见谅)
脚本代码如下:
import requests import itertools from bs4 import BeautifulSoup from csv import writer from random import randint from time import sleep from datetime import date url = "https://www.propertypal.com/property-for-sale/northern-ireland/page-" headers = { 'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/106.0.0.0 Safari/537.36'} filename = date.today().strftime("ni-listings-%Y-%m-%d.csv") with open(filename, 'w', encoding='utf8', newline='') as f: thewriter = writer(f) header = ['Address', 'Price'] thewriter.writerow(header) # for page in range(1, 3): for page in itertools.count(1): req = requests.get(f"{url}{page}", headers=headers) soup = BeautifulSoup(req.content, 'html.parser') for li in soup.find_all('li', class_="pp-property-box"): title = li.find('h2').text price = li.find('p', class_="pp-property-price").text info = [title, price] thewriter.writerow(info) sleep(randint(1, 5)) # this script scrapes all pages and records all listings and their prices in daily csv
核心优化方案
1. 处理HTTP状态码,避免无效爬取
当前脚本未判断请求是否成功,遇到302跳转直接解析内容必然出错。每次请求后先检查状态码,针对性处理:
req = session.get(f"{url}{page}", headers=headers) # 检查状态码,非200则处理 if req.status_code != 200: if req.status_code == 302: print(f"页面{page}跳转,尝试跟进") req = session.get(req.headers['Location'], headers=headers) else: print(f"请求页面{page}失败,状态码:{req.status_code}") # 延长等待时间后重试当前页 sleep(randint(5, 10)) continue # 跳过后续解析,重新请求当前页
2. 丰富请求头,模拟真实浏览器
仅添加User-Agent不足以规避反爬,补充更多浏览器常用头信息:
headers = { 'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/106.0.0.0 Safari/537.36', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8', 'Accept-Language': 'en-US,en;q=0.5', 'Accept-Encoding': 'gzip, deflate, br', 'Connection': 'keep-alive', 'Upgrade-Insecure-Requests': '1' }
3. 增加重试机制,应对临时限制
用requests的会话和重试适配器,处理网络波动或临时封禁:
from requests.adapters import HTTPAdapter from urllib3.util.retry import Retry # 创建会话并设置重试策略 session = requests.Session() retry = Retry( total=3, # 总重试次数 backoff_factor=1, # 重试间隔按1、2、4秒递增 status_forcelist=[429, 500, 502, 503, 504] # 需要重试的异常状态码 ) adapter = HTTPAdapter(max_retries=retry) session.mount('https://', adapter) session.mount('http://', adapter) # 后续用session.get代替requests.get req = session.get(f"{url}{page}", headers=headers)
4. 动态调整请求间隔,贴近人类行为
替换固定随机间隔,用正态分布生成更自然的等待时间:
from random import gauss # 生成均值3秒、标准差1秒的等待时间,确保不小于1秒 wait_time = max(1, gauss(3, 1)) sleep(wait_time)
5. 设置终止条件,避免无限循环
当前用itertools.count(1)无限循环,当页面无房源时及时停止:
listings = soup.find_all('li', class_="pp-property-box") if not listings: print(f"页面{page}无房源,爬取结束") break for li in listings: # 解析逻辑不变
6. 优化文件写入,减少IO开销
每次解析一条就写入一次效率低,先缓存一批数据再批量写入:
batch_data = [] batch_size = 50 # 每50条写入一次 for li in listings: title = li.find('h2').text.strip() # 清理多余空格换行 price = li.find('p', class_="pp-property-price").text.strip() batch_data.append([title, price]) # 达到批量大小就写入 if len(batch_data) >= batch_size: thewriter.writerows(batch_data) batch_data = [] # 写入剩余数据 if batch_data: thewriter.writerows(batch_data)
7. 遵守网站robots.txt规则
先查看目标网站的robots.txt文件,确认允许爬取的路径和频率,避免违反网站规则。
8. 代理IP(可选)
如果频繁被封禁,可使用代理IP分散请求来源,注意选择可靠的代理服务:
proxies = { 'http': 'http://your-proxy-ip:port', 'https': 'https://your-proxy-ip:port' } # 请求时添加proxies参数 req = session.get(f"{url}{page}", headers=headers, proxies=proxies)
内容的提问来源于stack exchange,提问作者cts
相关产品推荐
相关产品推荐

