如何结合Scrapy与Splash抓取JS页面及解决规则失效问题
Got it, let's break down why your rules aren't working and how to fix this properly.
The Core Problem
When you override start_requests and set the callback directly to first_tier, you're skipping the CrawlSpider's built-in parse method—which is exactly what processes your rules to extract and follow links. Additionally, even if the rules did run, they'd generate regular Request objects instead of SplashRequests, so JavaScript-rendered content still wouldn't be handled.
Step-by-Step Fix
1. Let CrawlSpider Handle Rules via Its Default Parse
Modify your start_requests to use self.parse (CrawlSpider's built-in callback) instead of first_tier. This ensures your rules are triggered:
def start_requests(self): for url in self.start_urls: yield SplashRequest( url, callback=self.parse, # Use CrawlSpider's parse to trigger rules endpoint='render.html', args={'wait': 3.5}, )
2. Make Rules Generate SplashRequests
Add a helper method to convert regular requests from your rules into SplashRequests, then attach it to each rule using the process_request parameter:
def _create_splash_request(self, request, response): # Convert regular Request to SplashRequest return SplashRequest( request.url, endpoint='render.html', args={'wait': 3.5}, meta=request.meta, # Preserve meta data from original request )
Update your rules to use this helper:
rules = ( Rule( LinkExtractor(allow=('catalogue\/category', ), deny=('subsection\.php', )), process_request='_create_splash_request' # Apply Splash to this rule's requests ), Rule( LinkExtractor(allow=('catalogue', ), deny=('catalogue\/category')), callback='first_tier', process_request='_create_splash_request' # Apply Splash here too ), )
3. Full Modified Spider Code
Here's the complete working version of your spider:
from scrapy.spiders import CrawlSpider, Rule from scrapy.linkextractors import LinkExtractor from scrapy_splash import SplashRequest class MySpider(CrawlSpider): name = 'booki' start_urls = [ 'https://worldmap.com/listings/in/united-states/', ] rules = ( Rule( LinkExtractor(allow=('catalogue\/category', ), deny=('subsection\.php', )), process_request='_create_splash_request' ), Rule( LinkExtractor(allow=('catalogue', ), deny=('catalogue\/category')), callback='first_tier', process_request='_create_splash_request' ), ) custom_settings = { 'SPLASH_URL': 'http://localhost:8050', 'DOWNLOADER_MIDDLEWARES': { 'scrapy_splash.SplashCookiesMiddleware': 723, 'scrapy_splash.SplashMiddleware': 725, 'scrapy.downloadermiddlewares.httpcompression.HttpCompressionMiddleware': 810, }, 'SPIDER_MIDDLEWARES': { 'scrapy_splash.SplashDeduplicateArgsMiddleware': 100, }, 'DUPEFILTER_CLASS': 'scrapy_splash.SplashAwareDupeFilter', 'DOWNLOAD_DELAY': 8, 'ITEM_PIPELINES': { 'bookstoscrap.pipelines.BookstoscrapPipeline': 300, } } def start_requests(self): for url in self.start_urls: yield SplashRequest( url, callback=self.parse, endpoint='render.html', args={'wait': 3.5}, ) def _create_splash_request(self, request, response): return SplashRequest( request.url, endpoint='render.html', args={'wait': 3.5}, meta=request.meta, ) def first_tier(self, response): # Your existing parsing logic here pass
Important Notes
- Never override
parsein a CrawlSpider: This method is reserved for processing rules—overwriting it will break your rule-based crawling. - Verify Splash is running: Make sure your Splash server is active at
http://localhost:8050before starting the spider. - Adjust wait time: The
wait: 3.5argument gives JavaScript time to render; tweak this if content still isn't loading properly.
内容的提问来源于stack exchange,提问作者Tajs

