如何用Python自动抓取Zillow房产Zestimate历史价格数据?
自动化批量抓取Zillow Zestimate历史数据方案指导
核心思路
你手动从Network Tab拿到的是Zillow的后端API接口数据,自动化的核心就是直接调用这些API,不需要渲染整个页面,这样效率更高也更稳定。步骤拆解下来就是:提取房产唯一标识(zpid)→构造API请求→批量获取数据→复用你的解析代码生成DataFrame。
具体步骤&代码实现
1. 从房产URL提取zpid
Zillow的每个房产都有唯一的zpid,从URL里提取这个ID是构造API请求的关键,用正则就能搞定:
import re def extract_zpid(zillow_url): # 适配两种常见Zillow URL格式:带_zpid后缀或直接以数字结尾 zpid_match = re.search(r'(\d+)(?:_zpid)?/', zillow_url) return zpid_match.group(1) if zpid_match else None
2. 构造API请求(复用你手动抓的接口)
把你在Network Tab里看到的API URL、请求头复制过来,模拟真实浏览器请求:
import requests import pandas as pd import time # 替换成你手动抓的请求头,重点保留User-Agent、Cookie、Referer REQUEST_HEADERS = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Cookie': '你的Cookie内容', 'Referer': 'https://www.zillow.com/' } def fetch_zestimate_data(zpid): # 替换成你手动抓的API路径,比如类似下面的格式(以实际为准) api_endpoint = f'https://www.zillow.com/async-create-zestimate-chart?zpid={zpid}' try: # 加延迟避免触发反爬 time.sleep(2) response = requests.get(api_endpoint, headers=REQUEST_HEADERS) response.raise_for_status() # 捕获请求错误 return response.json() except Exception as e: print(f"获取zpid {zpid} 数据失败: {str(e)}") return None
3. 批量处理&复用你的解析代码
把现有解析JSON生成DataFrame的逻辑套进去,批量遍历URL列表:
# 假设你已有的解析函数是这个(可根据你的实际代码调整) def parse_zestimate_json(json_data): if 'history' not in json_data: return pd.DataFrame() # 提取时间和价格字段 history_records = json_data['history'] date_list = [item['date'] for item in history_records] price_list = [item['value'] for item in history_records] # 返回带房产标识的DataFrame return pd.DataFrame({ 'Date': date_list, 'Zestimate Price': price_list }) def batch_process(url_list): all_property_data = [] for url in url_list: zpid = extract_zpid(url) if not zpid: print(f"无法从URL提取zpid: {url}") continue json_data = fetch_zestimate_data(zpid) if not json_data: continue property_df = parse_zestimate_json(json_data) property_df['Property URL'] = url # 标记数据来源URL all_property_data.append(property_df) print(f"完成抓取: {url}") # 合并所有数据 final_df = pd.concat(all_property_data, ignore_index=True) return final_df # 示例调用 your_url_list = [ 'https://www.zillow.com/homedetails/xxxxxxx/', # 你的57个URL... ] result_df = batch_process(your_url_list) result_df.to_csv('zillow_zestimate_10y_history.csv', index=False)
关键注意事项
- Cookie维护:Zillow的Cookie会过期,建议定期从浏览器复制更新,或者用
requests.Session()维持登录状态(需要手动登录一次获取session)。 - 反爬规避:Zillow对请求频率敏感,一定要加延迟(比如2-5秒/请求),如果后续处理数千个URL,建议搭配代理IP池分散请求。
- API兼容性:Zillow可能随时调整API接口,一旦发现抓取失败,先去Network Tab重新确认最新的API路径和参数。
- 数据完整性:部分房产的Zestimate历史可能不足10年,解析时要处理空数据或缺失字段的情况。
内容的提问来源于stack exchange,提问作者talha asif
相关产品推荐
相关产品推荐

