You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取Zillow旧金山房源页面仅获取少量价格数据的技术求助

解决Zillow房源数据爬取仅返回少量结果的问题

我来帮你拆解问题根源,再给出两种可行的解决办法:

问题本质

你用requests+BeautifulSoup只能拿到前9条数据,核心原因是Zillow的房源列表是动态加载的——初始页面只渲染了第一屏的房源,剩下的内容需要用户滚动页面后,才会通过AJAX请求加载更多数据。静态HTTP请求只能获取到页面初始渲染的HTML,自然拿不到后续加载的房源。

解决方案

思路1:用Selenium模拟浏览器滚动加载全部内容

Selenium可以模拟真实浏览器的操作,包括滚动页面、等待动态内容加载,完美适配这类动态渲染的网站。

步骤&代码示例:

  1. 先安装Selenium:
pip install selenium
  1. 下载对应浏览器的驱动(比如Chrome的ChromeDriver,要和你的浏览器版本匹配),放到项目目录或者系统PATH中。
  2. 编写爬取代码:
from selenium import webdriver
from bs4 import BeautifulSoup
import time

url = "https://www.zillow.com/homes/San-Francisco,-CA_rb/"
# 初始化Chrome浏览器(换成你常用的浏览器也可以,比如Firefox)
driver = webdriver.Chrome()
driver.get(url)

# 模拟滚动页面,直到所有房源加载完成
last_scroll_height = driver.execute_script("return document.body.scrollHeight")
while True:
    # 滚动到当前页面底部
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    # 等待3秒让内容加载(可根据网络情况调整)
    time.sleep(3)
    # 计算新的页面高度,判断是否已经加载完所有内容
    new_scroll_height = driver.execute_script("return document.body.scrollHeight")
    if new_scroll_height == last_scroll_height:
        break
    last_scroll_height = new_scroll_height

# 现在页面已加载全部房源,解析HTML并关闭浏览器
soup = BeautifulSoup(driver.page_source, "html.parser")
driver.quit()

# 提取房源价格(注意Zillow的元素属性可能会更新,需验证)
price_elements = soup.find_all("span", attrs={"data-test": "property-card-price"})
prices_list = [price.get_text().strip() for price in price_elements]
print(prices_list)

思路2:直接调用Zillow的内部API(更高效稳定)

通过浏览器开发者工具抓包,能发现Zillow滚动加载时会请求内部API获取结构化的房源数据。直接调用这个API可以跳过页面渲染步骤,效率更高,还能拿到更规整的JSON格式数据。

代码示例:

import requests

headers = {
    "Accept-Language": "en-GB,en;q=0.5",
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:104.0) Gecko/20100101 Firefox/104.0",
    "Referer": "https://www.zillow.com/homes/San-Francisco,-CA_rb/",
    # 这里需要填写你浏览器中的Zillow Cookie,可从开发者工具的网络请求中复制
    "Cookie": "你的Zillow Cookie内容"
}

# API接口及参数(注意参数可能随Zillow更新变动,需自行抓包验证)
api_url = "https://www.zillow.com/search/GetSearchPageState.htm"
params = {
    "searchQueryState": '{"pagination":{},"usersSearchTerm":"San Francisco, CA","mapBounds":{"west":-122.55178699316406,"east":-122.34425300683594,"south":37.69926902923061,"north":37.84028569417633},"regionSelection":[{"regionId":20330,"regionType":6}],"isMapVisible":true,"filterState":{"sortSelection":{"value":"globalrelevanceex"},"isAllHomes":{"value":true}},"isListVisible":true}',
    "wants": '{"cat1":["listResults","mapResults"],"cat2":["total"]}',
    "requestId": "3"
}

response = requests.get(api_url, headers=headers, params=params)
data = response.json()

# 提取所有房源价格
prices_list = []
# 提取列表视图的房源价格
for item in data["cat1"]["searchResults"]["listResults"]:
    prices_list.append(item["price"])
# 提取地图视图的房源价格(如果需要)
for item in data["cat1"]["searchResults"]["mapResults"]:
    prices_list.append(item["price"])

print(prices_list)

注意:Zillow有反爬机制,API参数和Cookie可能随时变动,需要你定期通过浏览器抓包更新;同时请求频率不要过高,避免IP被封禁。

内容的提问来源于stack exchange,提问作者salman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.27 14:42:39