使用Scrapy Selenium时如何对初始首页请求应用SeleniumRequest
初始首页使用SeleniumRequest的实现方法
核心逻辑是重写Scrapy爬虫默认的start_requests方法,把初始的普通请求替换为SeleniumRequest即可,原有详情页的逻辑不需要改动。
修改后的完整代码
import scrapy from scrapy_selenium import SeleniumRequest class ZoosSpider(scrapy.Spider): name = 'zoos' allowed_domains = ['www.tripadvisor.co.uk'] # 可保留start_urls配置,方便后续调整初始入口 start_urls = [ "https://www.tripadvisor.co.uk/Attractions-g186216-Activities-c53-a_allAttractions.true-United_Kingdom.html" ] existList = [] # 重写start_requests方法,替换初始请求为SeleniumRequest def start_requests(self): for url in self.start_urls: yield SeleniumRequest( url=url, wait_time=10, # 可根据首页加载速度调整等待时长 callback=self.parse ) def parse(self, response): tmpSEC = response.xpath("//section[@data-automation='AppPresentation_SingleFlexCardSection']") for elem in tmpSEC: link = response.urljoin(elem.xpath(".//a/@href").get()) yield SeleniumRequest( url=link, wait_time= 10, callback=self.parseDetails) def parseDetails(self, response): tmpName = response.xpath("//h1[@data-automation='mainH1']/text()").get() tmpLink = response.xpath("//div[@class='Lvkmj']/a/@href").getall() tmpURL = tmpTelnr = tmpMail = "N/A" yield { "Name": tmpName, "URL": tmpURL, }
关键说明
- 重写
start_requests方法后,Scrapy会优先使用该方法生成的请求,原有start_urls配置可以保留,方便后续批量调整初始爬取入口 - 若首页需要等待指定元素加载完成再返回响应,可以搭配
wait_until参数使用,避免拿到未渲染完成的页面,示例代码如下:
# 需先导入对应依赖 from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By yield SeleniumRequest( url=url, wait_time=10, # 等待列表卡片元素加载完成后再返回响应 wait_until=EC.presence_of_element_located((By.XPATH, "//section[@data-automation='AppPresentation_SingleFlexCardSection']")), callback=self.parse )
- 操作前请确认
settings.py中已经正确配置了scrapy_selenium的相关参数,包括下载中间件注册、浏览器驱动路径、无头模式等基础配置。
内容的提问来源于stack exchange,提问作者Rapid1898
相关产品推荐
相关产品推荐

