You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

BeautifulSoup无法获取Zillow页面全部地址,仅返回9条结果求助

解决Zillow地址爬取不全的问题

问题根源

Zillow的房源列表采用动态加载机制:初始HTTP请求仅返回页面可见区域的9条房源数据,剩余房源需要用户滚动页面到下方时,才会通过AJAX请求加载并渲染到页面中。直接用requests获取静态HTML只能拿到初始加载的内容,因此只能提取到9条地址。

解决方案:用Selenium模拟浏览器动态加载

Selenium可以模拟真实浏览器的操作,包括滚动页面触发加载,等所有内容渲染完成后再解析HTML。

代码实现

from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
import time

# 配置Chrome选项
chrome_options = Options()
chrome_options.add_argument("--headless=new")  # 无头模式,不显示浏览器窗口
chrome_options.add_argument("user-agent=Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_5) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/84.0.4147.125 Safari/537.36")
chrome_options.add_argument("accept-language=en-GB,en-US;q=0.9,en;q=0.8")

# 初始化浏览器
driver = webdriver.Chrome(options=chrome_options)
driver.get("https://www.zillow.com/homes/San-Francisco,-CA_rb/")

# 模拟滚动页面,加载全部内容
last_height = driver.execute_script("return document.body.scrollHeight")
while True:
    # 滚动到页面底部
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    # 等待加载完成
    time.sleep(3)
    # 获取新的页面高度
    new_height = driver.execute_script("return document.body.scrollHeight")
    # 高度不再变化则停止滚动
    if new_height == last_height:
        break
    last_height = new_height

# 获取渲染后的完整HTML并关闭浏览器
data = driver.page_source
driver.quit()

# 解析HTML提取地址
soup = BeautifulSoup(data, "html.parser")
tag_address = soup.find_all('address')

for x in tag_address:
    print(x.get_text(strip=True))

注意事项

  • 需提前安装Selenium和对应浏览器的驱动(如ChromeDriver),确保驱动版本与浏览器版本匹配。
  • 可根据网络情况调整time.sleep(3)的时长,保证内容有足够加载时间。
  • Zillow有反爬机制,频繁请求可能被限制,建议添加请求间隔或使用代理。

备选方案:直接调用Zillow的API

如果不想用Selenium,可以抓包找到Zillow加载房源的API接口,直接请求API获取JSON格式的房源数据,这种方式效率更高。不过API接口可能包含签名或参数验证,需要自行分析请求参数。

内容的提问来源于stack exchange,提问作者Wojciech madejski

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 07:41:24