You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何结合Scrapy与Splash抓取JS页面及解决规则失效问题

Fixing Scrapy CrawlSpider Rules Ignored When Using Splash

Got it, let's break down why your rules aren't working and how to fix this properly.

The Core Problem

When you override start_requests and set the callback directly to first_tier, you're skipping the CrawlSpider's built-in parse method—which is exactly what processes your rules to extract and follow links. Additionally, even if the rules did run, they'd generate regular Request objects instead of SplashRequests, so JavaScript-rendered content still wouldn't be handled.

Step-by-Step Fix

1. Let CrawlSpider Handle Rules via Its Default Parse

Modify your start_requests to use self.parse (CrawlSpider's built-in callback) instead of first_tier. This ensures your rules are triggered:

def start_requests(self):
    for url in self.start_urls:
        yield SplashRequest(
            url,
            callback=self.parse,  # Use CrawlSpider's parse to trigger rules
            endpoint='render.html',
            args={'wait': 3.5},
        )

2. Make Rules Generate SplashRequests

Add a helper method to convert regular requests from your rules into SplashRequests, then attach it to each rule using the process_request parameter:

def _create_splash_request(self, request, response):
    # Convert regular Request to SplashRequest
    return SplashRequest(
        request.url,
        endpoint='render.html',
        args={'wait': 3.5},
        meta=request.meta,  # Preserve meta data from original request
    )

Update your rules to use this helper:

rules = (
    Rule(
        LinkExtractor(allow=('catalogue\/category', ), deny=('subsection\.php', )),
        process_request='_create_splash_request'  # Apply Splash to this rule's requests
    ),
    Rule(
        LinkExtractor(allow=('catalogue', ), deny=('catalogue\/category')),
        callback='first_tier',
        process_request='_create_splash_request'  # Apply Splash here too
    ),
)

3. Full Modified Spider Code

Here's the complete working version of your spider:

from scrapy.spiders import CrawlSpider, Rule
from scrapy.linkextractors import LinkExtractor
from scrapy_splash import SplashRequest

class MySpider(CrawlSpider):
    name = 'booki'
    start_urls = [
        'https://worldmap.com/listings/in/united-states/',
    ]
    rules = (
        Rule(
            LinkExtractor(allow=('catalogue\/category', ), deny=('subsection\.php', )),
            process_request='_create_splash_request'
        ),
        Rule(
            LinkExtractor(allow=('catalogue', ), deny=('catalogue\/category')),
            callback='first_tier',
            process_request='_create_splash_request'
        ),
    )
    custom_settings = {
        'SPLASH_URL': 'http://localhost:8050',
        'DOWNLOADER_MIDDLEWARES': {
            'scrapy_splash.SplashCookiesMiddleware': 723,
            'scrapy_splash.SplashMiddleware': 725,
            'scrapy.downloadermiddlewares.httpcompression.HttpCompressionMiddleware': 810,
        },
        'SPIDER_MIDDLEWARES': {
            'scrapy_splash.SplashDeduplicateArgsMiddleware': 100,
        },
        'DUPEFILTER_CLASS': 'scrapy_splash.SplashAwareDupeFilter',
        'DOWNLOAD_DELAY': 8,
        'ITEM_PIPELINES': {
            'bookstoscrap.pipelines.BookstoscrapPipeline': 300,
        }
    }

    def start_requests(self):
        for url in self.start_urls:
            yield SplashRequest(
                url,
                callback=self.parse,
                endpoint='render.html',
                args={'wait': 3.5},
            )

    def _create_splash_request(self, request, response):
        return SplashRequest(
            request.url,
            endpoint='render.html',
            args={'wait': 3.5},
            meta=request.meta,
        )

    def first_tier(self, response):
        # Your existing parsing logic here
        pass

Important Notes

  • Never override parse in a CrawlSpider: This method is reserved for processing rules—overwriting it will break your rule-based crawling.
  • Verify Splash is running: Make sure your Splash server is active at http://localhost:8050 before starting the spider.
  • Adjust wait time: The wait: 3.5 argument gives JavaScript time to render; tweak this if content still isn't loading properly.

内容的提问来源于stack exchange,提问作者Tajs

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 03:56:13