使用BeautifulSoup爬取Zillow旧金山房源页面仅获取少量价格数据的技术求助
解决Zillow房源数据爬取仅返回少量结果的问题
我来帮你拆解问题根源,再给出两种可行的解决办法:
问题本质
你用requests+BeautifulSoup只能拿到前9条数据,核心原因是Zillow的房源列表是动态加载的——初始页面只渲染了第一屏的房源,剩下的内容需要用户滚动页面后,才会通过AJAX请求加载更多数据。静态HTTP请求只能获取到页面初始渲染的HTML,自然拿不到后续加载的房源。
解决方案
思路1:用Selenium模拟浏览器滚动加载全部内容
Selenium可以模拟真实浏览器的操作,包括滚动页面、等待动态内容加载,完美适配这类动态渲染的网站。
步骤&代码示例:
- 先安装Selenium:
pip install selenium
- 下载对应浏览器的驱动(比如Chrome的ChromeDriver,要和你的浏览器版本匹配),放到项目目录或者系统PATH中。
- 编写爬取代码:
from selenium import webdriver from bs4 import BeautifulSoup import time url = "https://www.zillow.com/homes/San-Francisco,-CA_rb/" # 初始化Chrome浏览器(换成你常用的浏览器也可以,比如Firefox) driver = webdriver.Chrome() driver.get(url) # 模拟滚动页面,直到所有房源加载完成 last_scroll_height = driver.execute_script("return document.body.scrollHeight") while True: # 滚动到当前页面底部 driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") # 等待3秒让内容加载(可根据网络情况调整) time.sleep(3) # 计算新的页面高度,判断是否已经加载完所有内容 new_scroll_height = driver.execute_script("return document.body.scrollHeight") if new_scroll_height == last_scroll_height: break last_scroll_height = new_scroll_height # 现在页面已加载全部房源,解析HTML并关闭浏览器 soup = BeautifulSoup(driver.page_source, "html.parser") driver.quit() # 提取房源价格(注意Zillow的元素属性可能会更新,需验证) price_elements = soup.find_all("span", attrs={"data-test": "property-card-price"}) prices_list = [price.get_text().strip() for price in price_elements] print(prices_list)
思路2:直接调用Zillow的内部API(更高效稳定)
通过浏览器开发者工具抓包,能发现Zillow滚动加载时会请求内部API获取结构化的房源数据。直接调用这个API可以跳过页面渲染步骤,效率更高,还能拿到更规整的JSON格式数据。
代码示例:
import requests headers = { "Accept-Language": "en-GB,en;q=0.5", "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:104.0) Gecko/20100101 Firefox/104.0", "Referer": "https://www.zillow.com/homes/San-Francisco,-CA_rb/", # 这里需要填写你浏览器中的Zillow Cookie,可从开发者工具的网络请求中复制 "Cookie": "你的Zillow Cookie内容" } # API接口及参数(注意参数可能随Zillow更新变动,需自行抓包验证) api_url = "https://www.zillow.com/search/GetSearchPageState.htm" params = { "searchQueryState": '{"pagination":{},"usersSearchTerm":"San Francisco, CA","mapBounds":{"west":-122.55178699316406,"east":-122.34425300683594,"south":37.69926902923061,"north":37.84028569417633},"regionSelection":[{"regionId":20330,"regionType":6}],"isMapVisible":true,"filterState":{"sortSelection":{"value":"globalrelevanceex"},"isAllHomes":{"value":true}},"isListVisible":true}', "wants": '{"cat1":["listResults","mapResults"],"cat2":["total"]}', "requestId": "3" } response = requests.get(api_url, headers=headers, params=params) data = response.json() # 提取所有房源价格 prices_list = [] # 提取列表视图的房源价格 for item in data["cat1"]["searchResults"]["listResults"]: prices_list.append(item["price"]) # 提取地图视图的房源价格(如果需要) for item in data["cat1"]["searchResults"]["mapResults"]: prices_list.append(item["price"]) print(prices_list)
注意:Zillow有反爬机制,API参数和Cookie可能随时变动,需要你定期通过浏览器抓包更新;同时请求频率不要过高,避免IP被封禁。
内容的提问来源于stack exchange,提问作者salman
相关产品推荐
相关产品推荐

