Scrapy+Selenium爬取返回200但无数据,请求技术排查
问题排查:Scrapy+Selenium爬取Yapo.cl无数据返回
问题概述
使用Scrapy结合Selenium爬取Yapo.cl的首都大区租房信息,通过滚动页面加载动态广告链接,请求广告页面返回200状态码,但无法提取标题等数据,需要获取包括经纬度在内的完整广告信息,排查问题原因并解决。
原爬虫代码
import scrapy from selenium import webdriver from selenium.webdriver.chrome.options import Options from scrapy.selector import Selector from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time class YapoSpider(scrapy.Spider): name = 'yapo' allowed_domains = ['yapo.cl'] start_urls = ['https://www.yapo.cl/region-metropolitana/inmuebles/inmuebles/arrendar?tipo-inmueble=departamento,casa&pagina=1'] def __init__(self): chrome_options = Options() chrome_options.add_argument("--headless") self.driver = webdriver.Chrome(options= chrome_options) def parse(self, response): self.driver.get(response.url) # parse speed incremento = 50 velocidad = 0.5 # scroll height altura_total = self.driver.execute_script("return document.body.scrollHeight") for posicion in range(0, altura_total, incremento): # scroll self.driver.execute_script(f"window.scrollTo(0, {posicion});") time.sleep(velocidad) # bottom page self.driver.execute_script(f"window.scrollTo(0, {altura_total});") sel = Selector(text=self.driver.page_source) # Selector Scrapy. for href in sel.xpath("//a[contains(@class,'card inmo subcategory-1240 category-1000 has-cover is-visible')]/@href").extract(): url = response.urljoin(href) yield scrapy.Request(url, callback=self.parse_dir_contents) def parse_dir_contents(self, response): title = response.xpath("//h1[@class='my-2 title order-1 ng-star-inserted']/text()").extract_first() yield {'title': title} def closed(self): self.driver.quit()
核心问题分析
- 广告页未经过动态渲染:
parse_dir_contents直接使用Scrapy默认下载的静态响应,而Yapo.cl的广告详情页是Angular动态渲染的,静态HTML中不存在目标元素,导致XPath匹配失败。 - 元素定位器不稳定:原XPath依赖
ng-star-inserted这类Angular动态生成的class,网站更新或渲染逻辑变化会导致定位失效;列表页的链接定位class也可能存在动态变化问题。 - 滚动加载逻辑不完善:固定步长+固定等待时间的滚动方式,可能未触发所有广告的加载,部分链接未被渲染出来;且未判断页面是否加载完成就停止滚动。
- Headless模式被检测:默认的Headless模式特征明显,网站可能限制内容加载,导致页面数据不完整。
修复方案
1. 让广告页请求经过Selenium渲染
修改parse_dir_contents方法,使用已初始化的driver加载详情页,再用Scrapy Selector解析渲染后的页面:
def parse_dir_contents(self, response): self.driver.get(response.url) # 等待标题元素加载完成 WebDriverWait(self.driver, 10).until( EC.presence_of_element_located((By.XPATH, "//h1[contains(@class, 'title')]")) ) sel = Selector(text=self.driver.page_source) # 提取标题 title = sel.xpath("//h1[contains(@class, 'title')]/text()").extract_first().strip() # 提取经纬度:通常在页面的script标签或meta中,示例从script解析JSON script_data = sel.xpath("//script[contains(text(), 'lat') and contains(text(), 'lng')]/text()").extract_first() if script_data: # 假设数据是类似 {"lat":-33.45, "lng":-70.66} 的格式,需根据实际调整 import json try: data = json.loads(script_data) latitude = data.get('lat') longitude = data.get('lng') except: latitude = longitude = None else: latitude = longitude = None yield { 'title': title, 'latitude': latitude, 'longitude': longitude }
2. 优化滚动加载逻辑
替换原滚动循环为动态判断加载完成的逻辑,确保所有广告都被渲染:
def parse(self, response): self.driver.get(response.url) # 动态滚动加载所有内容 last_height = self.driver.execute_script("return document.body.scrollHeight") while True: self.driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") # 等待内容加载,可根据网络情况调整时间 time.sleep(1.5) new_height = self.driver.execute_script("return document.body.scrollHeight") # 高度不再变化说明加载完成 if new_height == last_height: break last_height = new_height sel = Selector(text=self.driver.page_source) # 优化链接定位器,减少动态class依赖 for href in sel.xpath("//a[contains(@class, 'card inmo has-cover')]/@href").extract(): url = response.urljoin(href) yield scrapy.Request(url, callback=self.parse_dir_contents)
3. 优化Headless模式参数
避免网站检测Headless浏览器,添加模拟正常浏览器的配置:
def __init__(self): chrome_options = Options() # 使用新版Headless模式,更接近正常浏览器 chrome_options.add_argument("--headless=new") # 禁用自动化检测特征 chrome_options.add_argument("--disable-blink-features=AutomationControlled") # 设置正常的User-Agent chrome_options.add_argument("--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36") chrome_options.add_experimental_option("excludeSwitches", ["enable-automation"]) chrome_options.add_experimental_option('useAutomationExtension', False) # 可选:禁用图片加载提升速度 chrome_options.add_argument("--blink-settings=imagesEnabled=false") self.driver = webdriver.Chrome(options=chrome_options)
4. 完善元素定位策略
避免依赖动态生成的class,优先使用元素的文本、属性或层级关系定位:
- 列表页链接:用
//a[contains(@class, 'card inmo')]替代包含动态编号的class - 详情页标题:用
//h1[contains(@class, 'title')]替代包含ng-star-inserted的定位器
内容的提问来源于stack exchange,提问作者warforterritory
相关产品推荐
相关产品推荐

