You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium及BeautifulSoup爬取Zillow房源链接无法返回全部元素如何解决

Zillow出租房源爬取不全问题解决方案

问题根因

  • 静态请求(BeautifulSoup方案):Zillow房源采用动态异步加载逻辑,首次返回的HTML仅包含首屏前9条左右数据,剩余房源需用户滚动页面触发后端接口请求后才会渲染到DOM中,直接通过requests请求页面只能拿到首屏少量数据。
  • Selenium原生滚动方案:原有逻辑直接将滚动容器拉到底部,无法触发Zillow懒加载的视口检测规则,中间区间的房源不会加载;同时未处理分页逻辑,单页加载到40条左右上限后没有跳转下一页,且Selenium的自动化特征容易被Zillow反爬系统识别,限制返回数据量。

解决方案

方案1:优化Selenium爬取逻辑

首先添加反爬配置隐藏自动化特征,再修改为分段滚动逻辑,同时适配分页跳转:

import time
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# 配置反爬参数
chrome_options = Options()
chrome_options.add_argument("--disable-blink-features=AutomationControlled")
chrome_options.add_experimental_option("excludeSwitches", ["enable-automation"])
chrome_options.add_experimental_option("useAutomationExtension", False)

driver = webdriver.Chrome(options=chrome_options)
# 隐藏webdriver标识
driver.execute_cdp_cmd("Page.addScriptToEvaluateOnNewDocument", {
    "source": "Object.defineProperty(navigator, 'webdriver', {get: () => undefined})"
})

ZILLOW_URL = "你的Zillow搜索链接"
driver.get(ZILLOW_URL)
time.sleep(3)

scroll_container = driver.find_element(By.ID, 'search-page-list-container')
last_height = driver.execute_script("return arguments[0].scrollHeight", scroll_container)
all_links = set() # 自动去重

while True:
    # 分段滚动半屏,触发懒加载
    driver.execute_script("arguments[0].scrollTop += arguments[0].offsetHeight/2", scroll_container)
    time.sleep(2)
    # 提取当前已加载的房源链接
    link_tags = driver.find_elements(By.CSS_SELECTOR, ".list-card-info a")
    for tag in link_tags:
        href = tag.get_attribute("href")
        if href:
            all_links.add(href)
    # 判断是否滚动到当前页底部
    new_height = driver.execute_script("return arguments[0].scrollHeight", scroll_container)
    if new_height == last_height:
        # 尝试点击下一页
        try:
            next_btn = WebDriverWait(driver, 5).until(EC.element_to_be_clickable((By.CSS_SELECTOR, 'a[title="Next page"]')))
            next_btn.click()
            time.sleep(3)
            last_height = driver.execute_script("return arguments[0].scrollHeight", scroll_container)
        except:
            # 没有下一页则退出循环
            break
    last_height = new_height

# 输出所有链接
print(list(all_links))
driver.quit()

方案2:直接调用后端接口爬取(效率更高)

打开浏览器F12开发者工具,切换到「网络」选项卡,筛选「Fetch/XHR」请求,滚动页面时可以找到搜索接口,请求返回为JSON格式,直接包含全量房源数据,无需解析HTML。接口请求参数中的pagination字段可控制分页,修改页码即可批量拉取所有房源。

注意:请求接口时需要完整复制浏览器中的请求头(包含cookie、user-agent、referer等字段),控制请求频率,必要时使用代理IP池避免被反爬封禁。

内容的提问来源于stack exchange,提问作者Tharu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 16:39:05