You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Scrapy(Python2)爬取加载缓慢的动态页面?

老哥,你这个问题我太熟了!核心原因是Scrapy本身不执行JavaScript——它爬的是服务器直接返回的初始HTML,也就是那个带加载GIF的页面,而后续的内容是浏览器跑JS之后才生成的。你加的time.sleep()和download_delay根本没用,因为它们只是在请求前后等一等,完全触发不了页面的JS渲染流程。

给你两个靠谱的解决方案,按需选:

方案1:直接抓AJAX接口(优先选,效率高)

先打开浏览器的开发者工具(F12),切换到Network面板,刷新页面看看有没有XHR/Fetch请求——这些就是页面加载内容的接口。直接用Scrapy请求这些接口,比模拟浏览器快多了。举个Python2的例子:

import scrapy

class QuotesSpider(scrapy.Spider):
    name = "quotes"
    
    def start_requests(self):
        # 把这里换成你找到的真实AJAX接口URL
        ajax_api_url = "https://example.com/api/load-quotes"
        yield scrapy.Request(url=ajax_api_url, callback=self.parse_ajax)
    
    def parse_ajax(self, response):
        # 接口一般返回JSON,直接解析就行
        data = response.json()
        for quote_item in data.get('quotes', []):
            yield {
                'text': quote_item.get('content'),
                'author': quote_item.get('author_name')
            }

方案2:用Scrapy+Selenium模拟浏览器(复杂页面必备)

如果页面的动态逻辑特别绕,找不到AJAX接口,那就只能用真实浏览器来渲染页面了。Selenium可以帮你调用Chrome/Firefox,等页面加载完再抓内容。

首先得装依赖(Python2要装兼容版本的Selenium):

pip install selenium==2.53.6

然后修改你的爬虫代码:

import scrapy
from selenium import webdriver
from scrapy.http import HtmlResponse
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By

class QuotesSpider(scrapy.Spider):
    name = "quotes"
    
    def __init__(self):
        # 这里要填你ChromeDriver的路径,记得下对应浏览器版本的驱动
        self.driver = webdriver.Chrome(executable_path='/your/path/to/chromedriver')
    
    def start_requests(self):
        urls = ['https://example.com']
        for url in urls:
            yield scrapy.Request(url=url, callback=self.parse)
    
    def parse(self, response):
        self.driver.get(response.url)
        
        # 别用time.sleep!用显式等待更靠谱——等目标元素出现再继续
        try:
            # 最多等10秒,直到.quote这个元素加载出来(换成你要抓的元素选择器)
            WebDriverWait(self.driver, 10).until(
                EC.presence_of_element_located((By.CSS_SELECTOR, '.quote'))
            )
        except Exception as e:
            self.logger.error(f"等待页面加载超时: {e}")
            return
        
        # 拿到渲染后的页面源码,转成Scrapy的Response对象,方便后续解析
        rendered_html = self.driver.page_source
        new_response = HtmlResponse(url=response.url, body=rendered_html, encoding='utf-8')
        
        # 现在就可以正常用CSS/XPath提取内容了
        for quote in new_response.css('.quote'):
            yield {
                'text': quote.css('.text::text').extract_first(),
                'author': quote.css('.author::text').extract_first()
            }
    
    def closed(self, reason):
        # 爬虫结束记得关浏览器,不然会留进程
        self.driver.quit()

小提示:

  • 尽量用显式等待替代time.sleep(),前者是等元素出现再继续,后者是硬等,容易浪费时间或者错过加载时机。
  • 如果可以的话,建议升级到Python3——Python2早就停更了,现在Scrapy的新版本、更方便的工具(比如Playwright)都只支持Python3,用起来舒服多了。

内容的提问来源于stack exchange,提问作者Codieroot

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 06:22:29