新手求助:Scrapy爬虫随机终止于部分页面
Hey there! Let's figure out why your Scrapy crawler is crashing mid-run while scraping r/india. Since your traceback got cut off, I'll cover the most common culprits for this kind of issue, along with actionable fixes:
1. Selector Failures on Unexpected Post Types
Reddit has different post formats (stickied posts, removed/deleted posts, etc.) that might not have the elements you're trying to extract (like upvote counts). If your code assumes every post has those elements, it'll throw an error when it hits an exception case.
Fix: Use safe extraction methods with fallbacks, or wrap your parsing logic in a try-except block to catch errors without crashing the whole crawler.
Example code snippet for your parse method:
def parse_post(self, response): try: # Use default values to handle missing elements upvotes = response.css('.score::text').extract_first(default='0') title = response.css('h1._eYtD2XCVieq6emjKBH3m::text').extract_first(default='No Title') post_url = response.url yield { 'title': title, 'upvotes': upvotes, 'url': post_url } except Exception as e: # Log the error instead of crashing self.logger.error(f"Failed to parse {response.url}: {str(e)}") # Optional: Yield an error entry to track problematic URLs yield {'error': str(e), 'url': response.url}
2. Reddit's Anti-Scraping Measures
Reddit actively blocks scrapers that don't mimic human browsing behavior. If you're hitting rate limits or sending requests without proper headers, the server might return malformed or empty responses, breaking your parser.
Fix: Tweak your settings.py to play nice with Reddit:
# Set a real user-agent to avoid being flagged as a bot USER_AGENT = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' # Add a delay between requests DOWNLOAD_DELAY = 2 # Enable auto-throttling to adjust speed based on server response AUTOTHROTTLE_ENABLED = True AUTOTHROTTLE_START_DELAY = 1 AUTOTHROTTLE_MAX_DELAY = 5
3. Broken Pagination Logic
Your recursive pagination might be generating invalid after parameters, leading to 404 pages or malformed content that your crawler can't process.
Fix: Double-check how you're extracting the next page token. Here's a robust way to handle Reddit pagination:
def parse(self, response): # Parse all posts on the current page for post_container in response.css('div._1poyrkZ7g36PawDueRza-J'): post_link = post_container.css('a._3jOxDPIQ0KaOWpzvSQo-1s::attr(href)').extract_first() if post_link: yield response.follow(post_link, self.parse_post) # Get the next page URL safely next_page = response.css('span.next-button a::attr(href)').extract_first() if next_page: # Only follow if the URL is valid yield response.follow(next_page, self.parse)
4. Get the Full Error Traceback
Since your original error log was cut off, the exact issue might be something specific (like a KeyError from a missing dictionary key). To get the full picture:
Fix: Run your crawler with verbose logging to capture the complete traceback:
scrapy crawl your_spider_name --logfile=scrapy.log -L DEBUG
Open scrapy.log after the crash—this will show you exactly which line of code is failing and why.
内容的提问来源于stack exchange,提问作者jay3shh

