You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何爬取使用模板引擎的网站?Scrapy+Selenium爬取问题求助

问题描述

尝试使用scrapy和selenium爬取网站,得到的结果为[ {{ certificant.FirstName }} {{ certificant.LastName }} ]。原以为是页面未加载完成,添加WebDriverWait等待按钮显示后再提取数据,结果依旧。怀疑该结果来自动态渲染的模板引擎,求有效爬取的解决方法。当前代码如下:

import scrapy

from scrapy import Request

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By


class PjFx110Spider(scrapy.Spider):
    name = "pj_fx110"

    ROOT_URL = 'https://aplanner.ca'

    start_urls = [
        ROOT_URL
    ]

    def __init__(self):
        options = Options()
#         options.add_argument("--headless")
        self.driver = webdriver.Chrome('./chromedriver', options=options)

    def start_requests(self):
        for url in self.start_urls:
            yield scrapy.Request(url=url, callback=self.parse)


    def parse(self, response):
        self.driver.get(response.url)
        WebDriverWait(self.driver, 3600).until(EC.presence_of_element_located((By.ID, 'btnShowResults')))

        lists = response.css('.list-group')
        name = lists.xpath('//*[@id="FPlist"]/div/ul[1]/li/span[1]/text()').extract()
        print(name, '---------lists----------')
解决方法

核心问题是你用Selenium驱动浏览器加载了页面,但依然在使用Scrapy原始的response对象提取数据——这个response是Scrapy直接请求得到的静态HTML,没有经过Selenium渲染,所以拿到的是模板占位符而非真实内容。

修正步骤

  1. 从Selenium的driver中获取渲染后的页面源码,转换成Scrapy的Selector对象再提取数据
  2. 优化等待条件,确保目标数据元素已渲染完成(而非仅等待按钮出现)

修改后的代码示例

import scrapy
from scrapy import Request
from scrapy.selector import Selector
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By

class PjFx110Spider(scrapy.Spider):
    name = "pj_fx110"
    ROOT_URL = 'https://aplanner.ca'
    start_urls = [ROOT_URL]

    def __init__(self):
        options = Options()
        # options.add_argument("--headless")
        self.driver = webdriver.Chrome('./chromedriver', options=options)

    def start_requests(self):
        for url in self.start_urls:
            yield scrapy.Request(url=url, callback=self.parse)

    def parse(self, response):
        self.driver.get(response.url)
        # 等待目标数据所在元素加载完成,替代仅等待按钮
        WebDriverWait(self.driver, 10).until(
            EC.presence_of_element_located((By.XPATH, '//*[@id="FPlist"]/div/ul[1]/li/span[1]'))
        )
        # 获取渲染后的页面源码,转换成Scrapy Selector
        rendered_html = self.driver.page_source
        sel = Selector(text=rendered_html)
        
        # 用转换后的Selector提取真实数据
        name = sel.xpath('//*[@id="FPlist"]/div/ul[1]/li/span[1]/text()').extract()
        print(name, '---------lists----------')

    def closed(self, reason):
        # 爬虫结束时关闭浏览器,避免资源占用
        self.driver.quit()

额外建议

  • 不要设置3600秒的超长等待时间,10-30秒足够,避免无意义的资源消耗
  • 可以尝试使用scrapy-selenium中间件,更优雅地集成Scrapy与Selenium,减少手动处理逻辑
  • 如果网站数据是通过AJAX接口加载的,直接抓取接口会比用Selenium渲染页面效率更高

内容的提问来源于stack exchange,提问作者Dora

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 18:37:14