You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

BeautifulSoup抓取Zillow内容少于Chrome源代码及人机验证处理问询

Zillow页面抓取异常问题解决方案

问题根因

你遇到的两个现象是关联的:requests默认发送的请求特征过于明显,直接被Zillow的反爬系统识别为爬虫,返回的是人机验证页面而非正常的房源详情页,所以解析得到的内容远少于浏览器正常访问的结果。

解决方案

1. 基础请求伪装

首先给请求添加符合浏览器特征的请求头,避免被直接识别为python-requests爬虫,修改后的代码如下:

from bs4 import BeautifulSoup
import requests

# 模拟Chrome浏览器的请求头
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36",
    "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8",
    "Accept-Language": "en-US,en;q=0.9",
    "Referer": "https://www.zillow.com/"
}

url = 'https://www.zillow.com/homedetails/49-Mountain-St-Hartford-CT-06106/58139903_zpid/'
html = requests.get(url, headers=headers)
# 先打印状态码和返回文本长度判断是否正常
print(f"响应状态码:{html.status_code}")
print(f"返回内容长度:{len(html.text)}")
bs = BeautifulSoup(html.text,"html.parser")

如果请求后状态码返回403、或者内容长度依然远小于浏览器源码长度,说明已经触发了更严格的反爬规则,需要进行下一步处理。

2. 动态内容渲染处理

Zillow的部分页面内容是通过JavaScript动态加载的,静态请求拿到的初始HTML本身就不包含这部分数据,需要用无头浏览器模拟真实浏览器的运行逻辑,等JS渲染完成后再提取内容,示例用Selenium实现:

from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By

chrome_options = Options()
# 不需要显示浏览器窗口可以开启无头模式,需要手动过验证可以注释该行
chrome_options.add_argument("--headless=new")
chrome_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36")
# 禁用指纹识别特征
chrome_options.add_argument("--disable-blink-features=AutomationControlled")
chrome_options.add_experimental_option("excludeSwitches", ["enable-automation"])
chrome_options.add_experimental_option("useAutomationExtension", False)

driver = webdriver.Chrome(options=chrome_options)
# 移除webdriver特征,避免被反爬识别
driver.execute_script("Object.defineProperty(navigator, 'webdriver', {get: () => undefined})")

url = 'https://www.zillow.com/homedetails/49-Mountain-St-Hartford-CT-06106/58139903_zpid/'
driver.get(url)
# 等待页面核心元素加载完成,比固定sleep更高效
try:
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CLASS_NAME, "ds-home-details-chip"))
    )
except:
    print("页面加载超时或遇到验证")
page_source = driver.page_source
bs = BeautifulSoup(page_source, "html.parser")
print(bs.body.prettify())
driver.quit()

3. 人机验证处理方案

  • 降低抓取频率:两次请求的间隔设置在3-5秒以上,不要短时间内高频批量请求,避免触发流量异常检测
  • 使用代理IP:如果本地IP已经被限制,可搭配代理IP池使用,每次请求切换不同IP,降低单个IP的请求压力
  • 复用验证状态:可以先在本地Chrome浏览器手动访问Zillow完成人机验证,之后导出浏览器的Cookie添加到请求头中,或者直接在无头浏览器中复用本地Chrome的用户数据目录,不需要重复验证
  • 大规模抓取优先使用Zillow官方提供的开发者接口,避免违反网站使用条款。

内容的提问来源于stack exchange,提问作者Jason

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 18:06:03