如何用BeautifulSoup从Zillow页面提取房产相关信息?
解决Zillow爬虫find/find_all无结果及数据提取问题
问题根源
Zillow搜索页面采用React动态渲染,大部分房源数据并非直接嵌入HTML标签,而是藏在页面的script标签内的JSON结构里,直接用find()/find_all()查找普通DOM元素自然拿不到结果。
具体实现步骤
- 提取页面内嵌的JSON数据
先定位包含房源数据的script标签(通常id为__NEXT_DATA__),解析其中的JSON内容。 - 从JSON中提取目标字段
遍历房源列表,分别提取地址、租金估值、房源链接、房价估值。 - 分类存入列表
将各类数据分别存入对应列表,处理可能的空值情况。
完整代码示例
import requests from bs4 import BeautifulSoup import json # 目标页面(替换为你实际请求的休斯顿搜索页URL) url = "https://www.zillow.com/houston-tx/rentals/" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } # 获取页面内容 response = requests.get(url, headers=headers) soup = BeautifulSoup(response.text, "html.parser") # 提取内嵌的JSON数据 script_tag = soup.find("script", id="__NEXT_DATA__") if not script_tag: print("未找到目标script标签") exit() # 解析JSON data = json.loads(script_tag.string) # 定位房源列表(路径可能随Zillow页面更新变化,需根据实际JSON结构调整) search_results = data["props"]["pageProps"]["searchPageState"]["cat1"]["searchResults"]["listResults"] # 初始化存储列表 addresses = [] rent_zestimates = [] property_links = [] zestimates = [] # 遍历提取数据 for item in search_results: # 地址 addr = item.get("address", "") addresses.append(addr) # Rent Zestimate(租金估值) rent_zest = item.get("rentZestimate") or item.get("price") # 部分房源直接显示价格而非估值 rent_zestimates.append(rent_zest if rent_zest else "无数据") # 房源链接(补全域名) link = "https://www.zillow.com" + item.get("detailUrl", "") property_links.append(link) # Zestimate/房价信息 zest = item.get("zestimate") or item.get("price") zestimates.append(zest if zest else "无数据") # 打印测试(取前3条数据验证) print("地址列表:", addresses[:3]) print("租金估值列表:", rent_zestimates[:3]) print("房源链接列表:", property_links[:3]) print("房价估值列表:", zestimates[:3])
注意事项
- 必须设置User-Agent:Zillow会拦截无标识的请求,务必添加浏览器UA模拟正常访问。
- JSON路径可能更新:Zillow的页面结构会不定期调整,若代码失效,需重新查看
__NEXT_DATA__里的JSON层级,修正提取路径。 - 反爬限制:频繁请求会被Zillow封禁IP,建议添加请求间隔(
time.sleep(1-3)),或使用代理分散请求来源。
内容的提问来源于stack exchange,提问作者Oliver
相关产品推荐
相关产品推荐

