为何无法爬取Zillow全部房源价格?Python爬虫问题求助
解决Zillow房源租金爬取不完整的问题
问题根源
Zillow采用动态内容加载机制:初始页面仅渲染少量房源,剩余内容会在用户滚动页面或停留时通过JavaScript异步加载。直接用requests获取的只是初始HTML,自然拿不到完整数据;若Selenium未处理好加载逻辑,也会出现同样问题。
修复方案(基于Selenium)
以下是能获取完整数据的代码,核心是模拟滚动加载、等待内容渲染,同时规避基础反爬:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.chrome.options import Options from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time from bs4 import BeautifulSoup URL_ZILLOW = r"https://www.zillow.com/homes/for_rent/?searchQueryState=%7B%22pagination%22%3A%7B%7D%2C%22mapBounds%22%3A%7B%22west%22%3A-123.4663871665039%2C%22east%22%3A-121.7744926352539%2C%22south%22%3A37.03952097286371%2C%22north%22%3A38.19687379258651%7D%2C%22isMapVisible%22%3Atrue%2C%22filterState%22%3A%7B%22price%22%3A%7B%22max%22%3A872627%7D%2C%22beds%22%3A%7B%22min%22%3A1%7D%2C%22fore%22%3A%7B%22value%22%3Afalse%7D%2C%22mp%22%3A%7B%22max%22%3A3000%7D%2C%22auc%22%3A%7B%22value%22%3Afalse%7D%2C%22nc%22%3A%7B%22value%22%3Afalse%7D%2C%22fr%22%3A%7B%22value%22%3Atrue%7D%2C%22fsbo%22%3A%7B%22value%22%3Afalse%7D%2C%22cmsn%22%3A%7B%22value%22%3Afalse%7D%2C%22fsba%22%3A%7B%22value%22%3Afalse%7D%7D%2C%22isListVisible%22%3Atrue%2C%22mapZoom%22%3A9%7D" # 配置Chrome选项,规避基础反爬 chrome_options = Options() chrome_options.add_argument("--disable-blink-features=AutomationControlled") chrome_options.add_experimental_option("excludeSwitches", ["enable-automation"]) chrome_options.add_experimental_option('useAutomationExtension', False) chrome_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36") driver = webdriver.Chrome(options=chrome_options) driver.get(URL_ZILLOW) # 模拟滚动加载,直到没有新内容出现 last_height = driver.execute_script("return document.body.scrollHeight") while True: # 滚动到页面底部 driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") # 等待内容加载 time.sleep(3) # 计算新的页面高度并与上次比较 new_height = driver.execute_script("return document.body.scrollHeight") if new_height == last_height: break last_height = new_height # 等待所有价格元素加载完成 WebDriverWait(driver, 10).until( EC.presence_of_all_elements_located((By.CSS_SELECTOR, ".gMDnGj span")) ) # 获取页面源码解析 soup = BeautifulSoup(driver.page_source, "html.parser") prices = soup.select(".gMDnGj span") for price in prices: print(price.text.strip()) driver.quit()
关键说明
- 动态加载处理:通过循环滚动页面+等待,触发Zillow的异步加载逻辑,拿到全部房源内容。
- 反爬规避:禁用Selenium的自动化标识,使用真实的User-Agent,避免被网站识别为爬虫。
- 等待机制:用
WebDriverWait确保目标元素完全加载后再解析,避免因元素未渲染导致的遗漏。
注意:Zillow的页面结构和反爬策略可能随时变化,如果选择器失效,需要重新在浏览器开发者工具中抓取最新的元素选择器。
内容的提问来源于stack exchange,提问作者Mr-Sepi0l
相关产品推荐
相关产品推荐

