Scrapy集成Selenium爬取房产站点出现Crawled 0 pages错误如何解决
问题排查与修复方案
核心问题点
- 爬虫未配置初始启动请求,Scrapy启动后无待执行任务直接结束
- 代码中重复定义了
parse方法,后定义的逻辑覆盖了先定义的内容,导致请求逻辑完全失效 - 请求头中
User-Agent字段名拼写错误,自定义UA不生效可能触发反爬 - Selenium驱动路径配置错误,
which()方法不需要套在已有的绝对路径外 - Chrome无头启动参数写法存在版本兼容问题
修正后的爬虫代码
import scrapy from scrapy_selenium import SeleniumRequest class HomesSpider(scrapy.Spider): name = 'homes' def start_requests(self): yield SeleniumRequest( url='https://www.vivanuncios.com.mx/s-venta-inmuebles/queretaro/page-2/v1c1097l1021p2', wait_time=3, callback=self.parse_result ) def parse_result(self, response): homes = response.xpath("//div[@class='viewport-contents']/div") for home in homes: yield{ 'price': home.xpath(".//span[@class='value wrapper']/span[@class='ad-price']/text()").get(), 'location': home.xpath(".//div[@class='tile-location one-liner']/b/text()").get(), 'description': home.xpath(".//div[@class='tile-desc one-liner']/a/text()").get(), 'bedrooms': home.xpath(".//div[@class='chiplets-inline-block re-bedroom']/text()").get(), 'm2': home.xpath(".//div[@class='chiplets-inline-block surface-area']/text()").get() }
需要调整的settings.py配置
- 修改请求头键名:
DEFAULT_REQUEST_HEADERS = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/95.0.4638.54 Safari/537.36' }
- 调整其他配置项:
# 先关闭robots协议校验避免被拦截 ROBOTSTXT_OBEY = False # SELENIUM配置部分 from shutil import which SELENIUM_DRIVER_NAME = 'chrome' # 直接传绝对路径,不需要套which SELENIUM_DRIVER_EXECUTABLE_PATH = "C:\\Users\\Cesal\\projects\\real_state\\chromedriver.exe" # 用新版Chrome兼容的无头参数,调试时可以先注释这行查看浏览器运行状态 SELENIUM_DRIVER_ARGUMENTS=['--headless=new']
调试建议
如果修改后依然没有返回数据,可以先注释掉selenium的无头参数,运行爬虫查看Chrome浏览器是否能正常打开目标页面,是否出现人机验证、访问拦截的情况,确认页面可正常加载后再开启无头模式运行。
内容的提问来源于stack exchange,提问作者cesalomx
相关产品推荐
相关产品推荐

