You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Scrapy抓取动态加载网站?含charliesmithrealty.net爬取需求

解决Scrapy结合Selenium抓取动态过滤搜索结果的问题

我仔细看了你提供的代码和遇到的问题,咱们一步步拆解问题并给出可行的修复方案:

你的代码里的核心问题

  • 过时的元素定位方法:find_element_by_xpath这类旧方法已经被Selenium弃用,必须改用find_element(By.XPATH, "xxx")的标准格式
  • Driver生命周期错误:你在循环完分页后先关闭了driver,再去获取page_source,这时候driver已经销毁,必然会抛出异常
  • 分页逻辑漏洞:你先一直点击下一页直到无法点击,这时候之前页面的房源数据已经被新页面覆盖,最终只能拿到最后一页的内容
  • 元素定位不够精准:部分XPath选择器没有考虑页面动态渲染的延迟,容易出现元素未找到的报错

修复后的完整代码

下面是调整后的代码,解决了上述所有问题,能稳定抓取每一页的搜索结果:

import scrapy
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from scrapy.selector import Selector


class RealSpider(scrapy.Spider):
    name = 'real'
    start_urls = ['https://charliesmithrealty.net']

    def __init__(self):
        # 初始化Chrome Driver,添加无头模式提升运行效率
        options = webdriver.ChromeOptions()
        options.add_argument('--headless=new')
        options.add_argument('--window-size=1920,1080')
        self.driver = webdriver.Chrome(options=options)
        self.wait = WebDriverWait(self.driver, 10)  # 显式等待替代time.sleep,更可靠

    def parse(self, response):
        self.driver.get(response.url)

        # 1. 输入并选择邮政编码
        location_input = self.wait.until(
            EC.presence_of_element_located((By.ID, 'bfg-location-input'))
        )
        location_input.send_keys("80023")
        # 等待下拉选项出现并点击匹配的邮政编码
        zip_option = self.wait.until(
            EC.element_to_be_clickable((By.XPATH, "//li[@lookup_field='zip_code' and contains(text(), '80023')]"))
        )
        zip_option.click()

        # 2. 设置价格区间
        # 点击最小价格下拉框
        min_price_dropdown = self.wait.until(
            EC.element_to_be_clickable((By.XPATH, "//div[contains(@class, 'bfg-dropdown-min')]//button"))
        )
        min_price_dropdown.click()
        min_price_option = self.wait.until(
            EC.element_to_be_clickable((By.XPATH, "//li[contains(@class, 'bfg-option-list-min')]//li[text()='$150K']"))
        )
        min_price_option.click()

        # 点击最大价格下拉框
        max_price_dropdown = self.wait.until(
            EC.element_to_be_clickable((By.XPATH, "//div[contains(@class, 'bfg-dropdown-max')]//button"))
        )
        max_price_dropdown.click()
        max_price_option = self.wait.until(
            EC.element_to_be_clickable((By.XPATH, "//li[contains(@class, 'bfg-option-list-max')]//li[text()='$600K']"))
        )
        max_price_option.click()

        # 3. 设置卧室和卫生间数量
        # 卧室下拉选择(自定义下拉框需用JS触发事件)
        beds_select = self.wait.until(
            EC.presence_of_element_located((By.ID, 'BedsDropdown'))
        )
        self.driver.execute_script("arguments[0].value = '4';", beds_select)
        self.driver.execute_script("arguments[0].dispatchEvent(new Event('change'));", beds_select)

        # 卫生间下拉选择
        baths_select = self.wait.until(
            EC.presence_of_element_located((By.ID, 'BathsDropdown'))
        )
        self.driver.execute_script("arguments[0].value = '3';", baths_select)
        self.driver.execute_script("arguments[0].dispatchEvent(new Event('change'));", baths_select)

        # 4. 执行搜索
        search_btn = self.wait.until(
            EC.element_to_be_clickable((By.XPATH, "//button[@aria-label='Submit Search']"))
        )
        search_btn.click()
        self.wait.until(
            EC.presence_of_element_located((By.ID, 'bfg-map-gallery'))
        )  # 等待搜索结果加载完成

        # 5. 分页抓取每一页数据
        while True:
            # 抓取当前页的房源URL
            html = self.driver.page_source
            resp = Selector(text=html)
            for item in resp.xpath("//div[@id='bfg-map-gallery']//mbb-galleryitem/div/a"):
                yield {
                    'property_url': item.xpath("./@href").get()
                }

            # 尝试点击下一页
            try:
                next_btn = self.wait.until(
                    EC.element_to_be_clickable((By.XPATH, "//a[@data-page='next' and not(@disabled)]"))
                )
                next_btn.click()
                # 等待新页面加载完成(通过判断旧元素失效实现)
                self.wait.until(
                    EC.staleness_of(resp.xpath("//div[@id='bfg-map-gallery']").get())
                )
            except:
                # 没有下一页时退出循环
                break

        # 最后关闭Driver
        self.driver.quit()

关键优化点说明

  • 用显式等待替代time.sleep:WebDriverWait会等待元素出现/可点击后再执行操作,避免因网络延迟导致的元素未加载问题,比固定等待时间更灵活可靠
  • 精准定位元素:优先用ID定位,XPath选择器添加更具体的条件,避免因页面结构小变化导致定位失败
  • 处理自定义下拉框:部分下拉框是前端自定义组件,直接设置value后需要触发change事件,才能让页面识别到选择操作
  • 分页实时抓取:每点击一次下一页就立即抓取当前页内容,避免数据被新页面覆盖
  • 无头模式运行:减少资源占用,更适合服务器或后台运行爬虫

关于API接口的补充建议

你提到没找到JSON接口,这类房产网站通常用React/Vue等前端框架渲染,数据可能通过GraphQL或加密接口传输。如果不想用Selenium,可以尝试:

  • 查看Network面板的Fetch/XHR标签,筛选graphql相关请求,很多这类网站会用GraphQL查询数据
  • 复制请求Headers中的authorization或cookie等认证信息,直接构造请求获取数据

不过如果API接口有反爬机制,Selenium仍然是最直接、最不易被拦截的解决方案。

内容的提问来源于stack exchange,提问作者Raisul Islam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 10:45:28