You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup无法抓取Trulia搜索结果页全部40条房源数据

问题:Trulia房源抓取仅返回前7条数据的原因及解决办法

问题描述

尝试抓取Trulia.com搜索结果页的房源数据(价格、卧室数、浴室数、地址),页面应包含40条房源,但无论是原BeautifulSoup脚本还是改用Selenium加滚动的版本,都仅返回前7条数据。两个月前代码可正常抓取全部数据,推测是懒加载或网站反爬机制变更导致。

原因分析

  • 动态类名失效:原代码中使用的Text__TextBase-sc-27a633b1-0-div这类class名是前端框架动态生成的,Trulia更新前端后,这些类名已变更,导致选择器只能匹配到初始加载的少量元素。
  • 懒加载逻辑变更:Trulia可能调整了懒加载触发条件,不再是单纯滚动到页面底部,而是需要房源卡片进入可视区域才会加载数据,原滚动到底部的方式无法触发所有内容加载。
  • 代理/反爬限制:ScraperAPI返回的页面可能被Trulia的反爬机制拦截,仅返回初始的7条数据,或者未启用JavaScript渲染模式,无法加载动态内容。

解决办法

1. 使用稳定的元素选择器

放弃依赖动态生成的class名,改用data-testid或房源卡片的稳定父容器来定位元素。Trulia的房源卡片通常有统一的data-testid="property-card"属性,可以先定位所有卡片,再从卡片内提取所需字段。

2. 优化Selenium滚动逻辑

改为逐个滚动到房源卡片的位置,确保每个卡片都进入可视区域,触发加载。同时使用显式等待代替固定time.sleep,提升稳定性。

3. 直接用Selenium访问(或开启ScraperAPI渲染模式)

如果使用ScraperAPI,需确保开启JavaScript渲染参数(render=true);或者直接用Selenium访问目标URL,避免代理带来的限制。

修改后的代码示例

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup
import time

url = "https://www.trulia.com/for_sale/Hartford,CT/1p_beds/1p_baths/1p_sqft/SINGLE-FAMILY_HOME_type/"

chrome_options = Options()
chrome_options.add_argument("--headless=new")  # 新版无头模式更接近正常浏览器
chrome_options.add_argument("--disable-gpu")
chrome_options.add_argument("--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36")

driver = webdriver.Chrome(options=chrome_options)
driver.get(url)

# 显式等待初始房源加载
wait = WebDriverWait(driver, 10)
wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, '[data-testid="property-card"]')))

# 滚动加载所有房源
last_count = 0
while True:
    # 获取当前已加载的房源数量
    current_cards = driver.find_elements(By.CSS_SELECTOR, '[data-testid="property-card"]')
    current_count = len(current_cards)
    
    if current_count == last_count:
        break  # 没有新房源加载,退出循环
    
    last_count = current_count
    
    # 滚动到最后一个房源的位置,触发加载
    driver.execute_script("arguments[0].scrollIntoView({behavior: 'smooth', block: 'center'});", current_cards[-1])
    time.sleep(1.5)  # 给加载留时间

# 解析页面内容
html = driver.page_source
soup = BeautifulSoup(html, 'html.parser')

# 提取所有房源数据
properties = []
cards = soup.select('[data-testid="property-card"]')
for card in cards:
    price = card.select_one('[data-testid="property-price"]').text if card.select_one('[data-testid="property-price"]') else None
    beds = card.select_one('[data-testid="property-beds"]').text if card.select_one('[data-testid="property-beds"]') else None
    baths = card.select_one('[data-testid="property-baths"]').text if card.select_one('[data-testid="property-baths"]') else None
    address = card.select_one('[data-testid="property-address"]').text if card.select_one('[data-testid="property-address"]') else None
    
    properties.append({
        "price": price,
        "beds": beds,
        "baths": baths,
        "address": address
    })

# 输出结果
print(f"共抓取到{len(properties)}条房源")
for prop in properties:
    print(prop)

driver.quit()

关键改进点

  • 使用data-testid="property-card"定位房源卡片,选择器更稳定。
  • 循环滚动到最后一个房源,确保每批新房源都进入视口触发加载。
  • 用显式等待代替固定延时,提升代码稳定性。
  • 设置真实的User-Agent,避免被反爬机制识别。

内容的提问来源于stack exchange,提问作者DmGawlNYC

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 03:48:20