You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

优化Python房产爬虫脚本:解决封禁、超时及302/CORS报错

爬虫优化请求:解决302/CORS问题并完成全量爬取

我写了一个基础Python脚本,用来爬取房产网站数据,把房源地址和价格存入CSV文件。目标房源有5000多条,但现在脚本爬了大概2000条就超时了,控制台还出现302和CORS policy错误。我已经加了sleep(randint(1, 5))设置随机请求间隔,但还需要更多优化办法,希望能高效完成全量爬取,同时尊重目标网站、降低资源负载。(刚接触Python和爬虫,有基础错误请见谅)

脚本代码如下:

import requests
import itertools
from bs4 import BeautifulSoup
from csv import writer
from random import randint
from time import sleep
from datetime import date


url = "https://www.propertypal.com/property-for-sale/northern-ireland/page-"
headers = {
    'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/106.0.0.0 Safari/537.36'}
filename = date.today().strftime("ni-listings-%Y-%m-%d.csv")

with open(filename, 'w', encoding='utf8', newline='') as f:
    thewriter = writer(f)
    header = ['Address', 'Price']
    thewriter.writerow(header)

    # for page in range(1, 3):
    for page in itertools.count(1):
        req = requests.get(f"{url}{page}", headers=headers)
        soup = BeautifulSoup(req.content, 'html.parser')

        for li in soup.find_all('li', class_="pp-property-box"):
            title = li.find('h2').text
            price = li.find('p', class_="pp-property-price").text

            info = [title, price]
            thewriter.writerow(info)

        sleep(randint(1, 5))

# this script scrapes all pages and records all listings and their prices in daily csv

核心优化方案

1. 处理HTTP状态码,避免无效爬取

当前脚本未判断请求是否成功,遇到302跳转直接解析内容必然出错。每次请求后先检查状态码,针对性处理:

req = session.get(f"{url}{page}", headers=headers)
# 检查状态码,非200则处理
if req.status_code != 200:
    if req.status_code == 302:
        print(f"页面{page}跳转,尝试跟进")
        req = session.get(req.headers['Location'], headers=headers)
    else:
        print(f"请求页面{page}失败,状态码:{req.status_code}")
        # 延长等待时间后重试当前页
        sleep(randint(5, 10))
        continue  # 跳过后续解析,重新请求当前页

2. 丰富请求头,模拟真实浏览器

仅添加User-Agent不足以规避反爬,补充更多浏览器常用头信息:

headers = {
    'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/106.0.0.0 Safari/537.36',
    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8',
    'Accept-Language': 'en-US,en;q=0.5',
    'Accept-Encoding': 'gzip, deflate, br',
    'Connection': 'keep-alive',
    'Upgrade-Insecure-Requests': '1'
}

3. 增加重试机制,应对临时限制

用requests的会话和重试适配器,处理网络波动或临时封禁:

from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry

# 创建会话并设置重试策略
session = requests.Session()
retry = Retry(
    total=3,  # 总重试次数
    backoff_factor=1,  # 重试间隔按1、2、4秒递增
    status_forcelist=[429, 500, 502, 503, 504]  # 需要重试的异常状态码
)
adapter = HTTPAdapter(max_retries=retry)
session.mount('https://', adapter)
session.mount('http://', adapter)

# 后续用session.get代替requests.get
req = session.get(f"{url}{page}", headers=headers)

4. 动态调整请求间隔,贴近人类行为

替换固定随机间隔,用正态分布生成更自然的等待时间:

from random import gauss

# 生成均值3秒、标准差1秒的等待时间,确保不小于1秒
wait_time = max(1, gauss(3, 1))
sleep(wait_time)

5. 设置终止条件,避免无限循环

当前用itertools.count(1)无限循环,当页面无房源时及时停止:

listings = soup.find_all('li', class_="pp-property-box")
if not listings:
    print(f"页面{page}无房源,爬取结束")
    break
for li in listings:
    # 解析逻辑不变

6. 优化文件写入,减少IO开销

每次解析一条就写入一次效率低,先缓存一批数据再批量写入:

batch_data = []
batch_size = 50  # 每50条写入一次

for li in listings:
    title = li.find('h2').text.strip()  # 清理多余空格换行
    price = li.find('p', class_="pp-property-price").text.strip()
    batch_data.append([title, price])
    # 达到批量大小就写入
    if len(batch_data) >= batch_size:
        thewriter.writerows(batch_data)
        batch_data = []
# 写入剩余数据
if batch_data:
    thewriter.writerows(batch_data)

7. 遵守网站robots.txt规则

先查看目标网站的robots.txt文件,确认允许爬取的路径和频率,避免违反网站规则。

8. 代理IP(可选)

如果频繁被封禁,可使用代理IP分散请求来源,注意选择可靠的代理服务:

proxies = {
    'http': 'http://your-proxy-ip:port',
    'https': 'https://your-proxy-ip:port'
}
# 请求时添加proxies参数
req = session.get(f"{url}{page}", headers=headers, proxies=proxies)

内容的提问来源于stack exchange,提问作者cts

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 12:05:37