You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬取房产网站中断求助:第2页中途停止至第7页

问题排查与修复方案

核心问题分析

  • 文件写入模式错误:每次调用transform时用'w'模式打开CSV,会覆盖之前爬取的所有数据,导致看起来像是只爬了部分内容
  • 未捕获异常导致循环中断:area和price字段没有加try-except处理,当某个房产条目缺失这些元素时,代码会直接抛出异常终止循环
  • 条件过滤导致数据丢失:只有当firmstr不为空时才写入行,跳过了所有没有经纪公司信息的房产条目
  • 无反爬延时:短时间内连续请求可能触发网站反爬机制,导致请求被拦截或返回不完整内容
  • 页码范围错误:range(1,7)仅会爬取1-6页,无法覆盖目标的第7页

修正后的完整代码

from bs4 import BeautifulSoup
import requests
import csv
import time

def extract(page):
    URL = f'https://www.point2homes.com/MX/Real-Estate-Listings.html?LocationGeoId=&LocationGeoAreaId=240589&Location=San%20Felipe,%20Baja%20California,%20Mexico&page={page}'
    headers = { 
        'Accept':'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8', 
        'User-Agent':'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/97.0.4692.71 Safari/537.36'
    }
    # 添加延时避免触发反爬
    time.sleep(2)
    r = requests.get(url=URL, headers=headers)
    # 检查请求是否成功
    if r.status_code != 200:
        print(f"请求第{page}页失败,状态码:{r.status_code}")
        return None
    soup = BeautifulSoup(r.content, 'html5lib')
    return soup

def transform(soup, is_first_page=False):
    if not soup:
        return
    listing = soup.findAll('article')
    # 首次写入用'w'创建文件,后续用'a'追加数据
    mode = 'w' if is_first_page else 'a'
    with open('housing.csv', mode, encoding='utf8', newline='') as f:
        thewriter = csv.writer(f)
        if is_first_page:
            header = ['Address', 'Beds', 'Baths', 'Size', 'Area', 'Acres', 'Price', 'Agent', 'Firm']
            thewriter.writerow(header)
        for ls in listing:
            # 地址提取
            try:
                address = ls.find('div', class_="address-container").text.replace('\n', "").strip()
            except:
                address = "Not Found"
            # 卧室数提取
            try:
                beds = ls.find('li', class_="ic-beds").text.replace('\n', "").strip()
            except:
                beds = "Data Not Logged"
            # 浴室数提取
            try:
                baths = ls.find('li', class_="ic-baths").text.replace('\n', "").strip()
            except:
                baths = "Data Not Logged"
            # 面积提取
            try:
                size = ls.find('li', class_="ic-sqft").text.replace('\n', "").strip()
            except:
                size = "Data Not Logged"
            # 占地提取
            try:
                acre = ls.find('li', class_="ic-lotsize").text.replace('\n', "").strip()
            except:
                acre = "Data Not Found"
            # 房产类型提取(新增异常捕获)
            try:
                area = ls.find('li', class_="property-type ic-proptype").text.replace('\n', "").strip()
            except:
                area = "Data Not Logged"
            # 价格提取(新增异常捕获)
            try:
                price = ls.find('span', class_="green").text.replace('\n', "").strip()
            except:
                price = "Data Not Logged"
            # 经纪人提取
            try:
                agent = ls.find('div', class_="agent-name").text.replace('\n', "").strip()
            except:
                agent = "Not Found"
            # 经纪公司提取(改用BeautifulSoup原生方法,替代字符串分割)
            try:
                firm = ls.find('div', class_="agent-company").text.replace('\n', "").strip()
            except:
                firm = "Not Found"
            
            # 无论字段是否为空,都写入数据行
            r_list = [address, beds, baths, size, area, acre, price, agent, firm]
            thewriter.writerow(r_list)
            print(r_list)

# 修正页码范围,range(1,8)覆盖1-7页
for i in range(1, 8):
    print(f'正在爬取第{i}页...')
    c = extract(i)
    transform(c, is_first_page=(i==1))

关键修复点说明

  1. 文件写入逻辑优化:区分首次创建文件和后续追加数据,避免覆盖之前爬取的内容
  2. 全面异常处理:给所有字段添加try-except,确保单个条目缺失数据时不会中断整个页面的爬取流程
  3. 取消数据过滤:移除仅当经纪公司存在才写入的限制,避免丢失无经纪公司信息的房产数据
  4. 优化数据提取方式:用BeautifulSoup原生的text方法提取经纪公司信息,替代易出错的字符串分割
  5. 反爬防护:添加请求延时,降低被网站反爬机制拦截的概率
  6. 请求状态校验:检查HTTP响应状态码,及时发现请求失败的情况
  7. 页码范围修正:调整range参数,确保覆盖1-7页的目标内容

内容的提问来源于stack exchange,提问作者Christopher Montoya

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 01:41:02