如何爬取使用模板引擎的网站?Scrapy+Selenium爬取问题求助
问题描述
尝试使用scrapy和selenium爬取网站,得到的结果为[ {{ certificant.FirstName }} {{ certificant.LastName }} ]。原以为是页面未加载完成,添加WebDriverWait等待按钮显示后再提取数据,结果依旧。怀疑该结果来自动态渲染的模板引擎,求有效爬取的解决方法。当前代码如下:
import scrapy from scrapy import Request from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By class PjFx110Spider(scrapy.Spider): name = "pj_fx110" ROOT_URL = 'https://aplanner.ca' start_urls = [ ROOT_URL ] def __init__(self): options = Options() # options.add_argument("--headless") self.driver = webdriver.Chrome('./chromedriver', options=options) def start_requests(self): for url in self.start_urls: yield scrapy.Request(url=url, callback=self.parse) def parse(self, response): self.driver.get(response.url) WebDriverWait(self.driver, 3600).until(EC.presence_of_element_located((By.ID, 'btnShowResults'))) lists = response.css('.list-group') name = lists.xpath('//*[@id="FPlist"]/div/ul[1]/li/span[1]/text()').extract() print(name, '---------lists----------')
解决方法
核心问题是你用Selenium驱动浏览器加载了页面,但依然在使用Scrapy原始的response对象提取数据——这个response是Scrapy直接请求得到的静态HTML,没有经过Selenium渲染,所以拿到的是模板占位符而非真实内容。
修正步骤
- 从Selenium的
driver中获取渲染后的页面源码,转换成Scrapy的Selector对象再提取数据 - 优化等待条件,确保目标数据元素已渲染完成(而非仅等待按钮出现)
修改后的代码示例
import scrapy from scrapy import Request from scrapy.selector import Selector from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By class PjFx110Spider(scrapy.Spider): name = "pj_fx110" ROOT_URL = 'https://aplanner.ca' start_urls = [ROOT_URL] def __init__(self): options = Options() # options.add_argument("--headless") self.driver = webdriver.Chrome('./chromedriver', options=options) def start_requests(self): for url in self.start_urls: yield scrapy.Request(url=url, callback=self.parse) def parse(self, response): self.driver.get(response.url) # 等待目标数据所在元素加载完成,替代仅等待按钮 WebDriverWait(self.driver, 10).until( EC.presence_of_element_located((By.XPATH, '//*[@id="FPlist"]/div/ul[1]/li/span[1]')) ) # 获取渲染后的页面源码,转换成Scrapy Selector rendered_html = self.driver.page_source sel = Selector(text=rendered_html) # 用转换后的Selector提取真实数据 name = sel.xpath('//*[@id="FPlist"]/div/ul[1]/li/span[1]/text()').extract() print(name, '---------lists----------') def closed(self, reason): # 爬虫结束时关闭浏览器,避免资源占用 self.driver.quit()
额外建议
- 不要设置3600秒的超长等待时间,10-30秒足够,避免无意义的资源消耗
- 可以尝试使用
scrapy-selenium中间件,更优雅地集成Scrapy与Selenium,减少手动处理逻辑 - 如果网站数据是通过AJAX接口加载的,直接抓取接口会比用Selenium渲染页面效率更高
内容的提问来源于stack exchange,提问作者Dora
相关产品推荐
相关产品推荐

