You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Zillow房价历史数据爬取遇Nonetype错误求助

解决Zillow爬取时出现的Nonetype AttributeError问题

问题根源

  1. 动态页面渲染:Zillow的房价历史数据是通过JavaScript异步加载的,requests仅能获取初始静态HTML,无法拿到后续动态生成的内容,导致目标元素不存在。
  2. 反爬拦截:Zillow会拦截无浏览器标识的请求,返回不完整页面,直接造成元素查找失败。
  3. 动态类名:你使用的类名(如hdp__sc-1j01zad-0 hGwlRq)是前端框架动态生成的,会随页面更新变化,依赖这类类名爬取稳定性极差。

解决方案

方案一:添加请求头模拟浏览器

先尝试给请求添加浏览器标识,获取完整页面:

import requests
from bs4 import BeautifulSoup

url = input('input url')
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
}

response = requests.get(url, headers=headers)
response.raise_for_status()  # 检查请求是否成功

soup = BeautifulSoup(response.content, 'html.parser')

# 通过文本定位价格历史区域,替代动态类名
price_history_section = soup.find('section', text=lambda t: t and 'Price History' in t)
if price_history_section:
    price_history_table = price_history_section.find_next('table')
    if price_history_table:
        # 跳过表头,遍历数据行
        rows = price_history_table.find_all('tr')[1:]
        for row in rows:
            cols = row.find_all('td')
            if len(cols) >= 2:
                date = cols[0].get_text(strip=True)
                price = cols[1].get_text(strip=True)
                print(f'{date}: {price}')
    else:
        print('未找到价格历史表格')
else:
    print('未找到价格历史区域')

方案二:使用Selenium获取动态渲染页面

如果方案一无效,说明内容完全依赖JS渲染,需用Selenium模拟浏览器:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time

url = input('input url')

# 初始化Chrome浏览器(需提前安装ChromeDriver并配置环境变量)
driver = webdriver.Chrome()
driver.get(url)

try:
    # 等待价格历史表格加载完成,最长等待10秒
    price_history_table = WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.XPATH, "//section[contains(., 'Price History')]//table"))
    )
    
    # 遍历数据行,跳过表头
    rows = price_history_table.find_elements(By.TAG_NAME, 'tr')[1:]
    for row in rows:
        cols = row.find_elements(By.TAG_NAME, 'td')
        date = cols[0].text.strip()
        price = cols[1].text.strip()
        print(f'{date}: {price}')
finally:
    # 关闭浏览器
    driver.quit()

注意事项

  • Zillow反爬机制严格,爬取时需添加请求延迟(如time.sleep(2)),避免频繁请求导致IP被封禁。
  • 定位元素时优先使用XPath、文本内容或固定data-*属性,避免依赖动态生成的类名。

内容的提问来源于stack exchange,提问作者Dataman20

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 08:42:51