You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy+Selenium爬取返回200但无数据,请求技术排查

问题排查:Scrapy+Selenium爬取Yapo.cl无数据返回

问题概述

使用Scrapy结合Selenium爬取Yapo.cl的首都大区租房信息,通过滚动页面加载动态广告链接,请求广告页面返回200状态码,但无法提取标题等数据,需要获取包括经纬度在内的完整广告信息,排查问题原因并解决。

原爬虫代码

import scrapy
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from scrapy.selector import Selector
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time

class YapoSpider(scrapy.Spider):
    name = 'yapo'
    allowed_domains = ['yapo.cl']
    start_urls = ['https://www.yapo.cl/region-metropolitana/inmuebles/inmuebles/arrendar?tipo-inmueble=departamento,casa&pagina=1']

    def __init__(self):
        chrome_options = Options()
        chrome_options.add_argument("--headless")
        self.driver = webdriver.Chrome(options= chrome_options)

    def parse(self, response):
        self.driver.get(response.url)

        # parse speed
        incremento = 50 
        velocidad = 0.5  

        # scroll height
        altura_total = self.driver.execute_script("return document.body.scrollHeight")

        for posicion in range(0, altura_total, incremento):
            # scroll
            self.driver.execute_script(f"window.scrollTo(0, {posicion});")
            time.sleep(velocidad)

        # bottom page
        self.driver.execute_script(f"window.scrollTo(0, {altura_total});")
        sel = Selector(text=self.driver.page_source)
        # Selector Scrapy.
        for href in sel.xpath("//a[contains(@class,'card inmo subcategory-1240 category-1000 has-cover is-visible')]/@href").extract():
            url = response.urljoin(href)
            yield scrapy.Request(url, callback=self.parse_dir_contents)

    def parse_dir_contents(self, response):
        title = response.xpath("//h1[@class='my-2 title order-1 ng-star-inserted']/text()").extract_first()
        yield {'title': title}

    def closed(self):
        self.driver.quit()

核心问题分析

  1. 广告页未经过动态渲染:parse_dir_contents直接使用Scrapy默认下载的静态响应,而Yapo.cl的广告详情页是Angular动态渲染的,静态HTML中不存在目标元素,导致XPath匹配失败。
  2. 元素定位器不稳定:原XPath依赖ng-star-inserted这类Angular动态生成的class,网站更新或渲染逻辑变化会导致定位失效;列表页的链接定位class也可能存在动态变化问题。
  3. 滚动加载逻辑不完善:固定步长+固定等待时间的滚动方式,可能未触发所有广告的加载,部分链接未被渲染出来;且未判断页面是否加载完成就停止滚动。
  4. Headless模式被检测:默认的Headless模式特征明显,网站可能限制内容加载,导致页面数据不完整。

修复方案

1. 让广告页请求经过Selenium渲染

修改parse_dir_contents方法,使用已初始化的driver加载详情页,再用Scrapy Selector解析渲染后的页面:

def parse_dir_contents(self, response):
    self.driver.get(response.url)
    # 等待标题元素加载完成
    WebDriverWait(self.driver, 10).until(
        EC.presence_of_element_located((By.XPATH, "//h1[contains(@class, 'title')]"))
    )
    sel = Selector(text=self.driver.page_source)
    # 提取标题
    title = sel.xpath("//h1[contains(@class, 'title')]/text()").extract_first().strip()
    # 提取经纬度:通常在页面的script标签或meta中,示例从script解析JSON
    script_data = sel.xpath("//script[contains(text(), 'lat') and contains(text(), 'lng')]/text()").extract_first()
    if script_data:
        # 假设数据是类似 {"lat":-33.45, "lng":-70.66} 的格式,需根据实际调整
        import json
        try:
            data = json.loads(script_data)
            latitude = data.get('lat')
            longitude = data.get('lng')
        except:
            latitude = longitude = None
    else:
        latitude = longitude = None
    yield {
        'title': title,
        'latitude': latitude,
        'longitude': longitude
    }

2. 优化滚动加载逻辑

替换原滚动循环为动态判断加载完成的逻辑,确保所有广告都被渲染:

def parse(self, response):
    self.driver.get(response.url)
    
    # 动态滚动加载所有内容
    last_height = self.driver.execute_script("return document.body.scrollHeight")
    while True:
        self.driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
        # 等待内容加载,可根据网络情况调整时间
        time.sleep(1.5)
        new_height = self.driver.execute_script("return document.body.scrollHeight")
        # 高度不再变化说明加载完成
        if new_height == last_height:
            break
        last_height = new_height
    
    sel = Selector(text=self.driver.page_source)
    # 优化链接定位器,减少动态class依赖
    for href in sel.xpath("//a[contains(@class, 'card inmo has-cover')]/@href").extract():
        url = response.urljoin(href)
        yield scrapy.Request(url, callback=self.parse_dir_contents)

3. 优化Headless模式参数

避免网站检测Headless浏览器,添加模拟正常浏览器的配置:

def __init__(self):
    chrome_options = Options()
    # 使用新版Headless模式,更接近正常浏览器
    chrome_options.add_argument("--headless=new")
    # 禁用自动化检测特征
    chrome_options.add_argument("--disable-blink-features=AutomationControlled")
    # 设置正常的User-Agent
    chrome_options.add_argument("--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36")
    chrome_options.add_experimental_option("excludeSwitches", ["enable-automation"])
    chrome_options.add_experimental_option('useAutomationExtension', False)
    # 可选:禁用图片加载提升速度
    chrome_options.add_argument("--blink-settings=imagesEnabled=false")
    self.driver = webdriver.Chrome(options=chrome_options)

4. 完善元素定位策略

避免依赖动态生成的class,优先使用元素的文本、属性或层级关系定位:

  • 列表页链接:用//a[contains(@class, 'card inmo')]替代包含动态编号的class
  • 详情页标题:用//h1[contains(@class, 'title')]替代包含ng-star-inserted的定位器

内容的提问来源于stack exchange,提问作者warforterritory

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 10:20:58