Python爬取房产网站中断求助:第2页中途停止至第7页
问题排查与修复方案
核心问题分析
- 文件写入模式错误:每次调用
transform时用'w'模式打开CSV,会覆盖之前爬取的所有数据,导致看起来像是只爬了部分内容 - 未捕获异常导致循环中断:
area和price字段没有加try-except处理,当某个房产条目缺失这些元素时,代码会直接抛出异常终止循环 - 条件过滤导致数据丢失:只有当
firmstr不为空时才写入行,跳过了所有没有经纪公司信息的房产条目 - 无反爬延时:短时间内连续请求可能触发网站反爬机制,导致请求被拦截或返回不完整内容
- 页码范围错误:
range(1,7)仅会爬取1-6页,无法覆盖目标的第7页
修正后的完整代码
from bs4 import BeautifulSoup import requests import csv import time def extract(page): URL = f'https://www.point2homes.com/MX/Real-Estate-Listings.html?LocationGeoId=&LocationGeoAreaId=240589&Location=San%20Felipe,%20Baja%20California,%20Mexico&page={page}' headers = { 'Accept':'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8', 'User-Agent':'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/97.0.4692.71 Safari/537.36' } # 添加延时避免触发反爬 time.sleep(2) r = requests.get(url=URL, headers=headers) # 检查请求是否成功 if r.status_code != 200: print(f"请求第{page}页失败,状态码:{r.status_code}") return None soup = BeautifulSoup(r.content, 'html5lib') return soup def transform(soup, is_first_page=False): if not soup: return listing = soup.findAll('article') # 首次写入用'w'创建文件,后续用'a'追加数据 mode = 'w' if is_first_page else 'a' with open('housing.csv', mode, encoding='utf8', newline='') as f: thewriter = csv.writer(f) if is_first_page: header = ['Address', 'Beds', 'Baths', 'Size', 'Area', 'Acres', 'Price', 'Agent', 'Firm'] thewriter.writerow(header) for ls in listing: # 地址提取 try: address = ls.find('div', class_="address-container").text.replace('\n', "").strip() except: address = "Not Found" # 卧室数提取 try: beds = ls.find('li', class_="ic-beds").text.replace('\n', "").strip() except: beds = "Data Not Logged" # 浴室数提取 try: baths = ls.find('li', class_="ic-baths").text.replace('\n', "").strip() except: baths = "Data Not Logged" # 面积提取 try: size = ls.find('li', class_="ic-sqft").text.replace('\n', "").strip() except: size = "Data Not Logged" # 占地提取 try: acre = ls.find('li', class_="ic-lotsize").text.replace('\n', "").strip() except: acre = "Data Not Found" # 房产类型提取(新增异常捕获) try: area = ls.find('li', class_="property-type ic-proptype").text.replace('\n', "").strip() except: area = "Data Not Logged" # 价格提取(新增异常捕获) try: price = ls.find('span', class_="green").text.replace('\n', "").strip() except: price = "Data Not Logged" # 经纪人提取 try: agent = ls.find('div', class_="agent-name").text.replace('\n', "").strip() except: agent = "Not Found" # 经纪公司提取(改用BeautifulSoup原生方法,替代字符串分割) try: firm = ls.find('div', class_="agent-company").text.replace('\n', "").strip() except: firm = "Not Found" # 无论字段是否为空,都写入数据行 r_list = [address, beds, baths, size, area, acre, price, agent, firm] thewriter.writerow(r_list) print(r_list) # 修正页码范围,range(1,8)覆盖1-7页 for i in range(1, 8): print(f'正在爬取第{i}页...') c = extract(i) transform(c, is_first_page=(i==1))
关键修复点说明
- 文件写入逻辑优化:区分首次创建文件和后续追加数据,避免覆盖之前爬取的内容
- 全面异常处理:给所有字段添加
try-except,确保单个条目缺失数据时不会中断整个页面的爬取流程 - 取消数据过滤:移除仅当经纪公司存在才写入的限制,避免丢失无经纪公司信息的房产数据
- 优化数据提取方式:用BeautifulSoup原生的
text方法提取经纪公司信息,替代易出错的字符串分割 - 反爬防护:添加请求延时,降低被网站反爬机制拦截的概率
- 请求状态校验:检查HTTP响应状态码,及时发现请求失败的情况
- 页码范围修正:调整
range参数,确保覆盖1-7页的目标内容
内容的提问来源于stack exchange,提问作者Christopher Montoya
相关产品推荐
相关产品推荐

