You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

新手求助:Scrapy爬虫随机终止于部分页面

Troubleshooting Your Scrapy Reddit Crawler Crash

Hey there! Let's figure out why your Scrapy crawler is crashing mid-run while scraping r/india. Since your traceback got cut off, I'll cover the most common culprits for this kind of issue, along with actionable fixes:

1. Selector Failures on Unexpected Post Types

Reddit has different post formats (stickied posts, removed/deleted posts, etc.) that might not have the elements you're trying to extract (like upvote counts). If your code assumes every post has those elements, it'll throw an error when it hits an exception case.

Fix: Use safe extraction methods with fallbacks, or wrap your parsing logic in a try-except block to catch errors without crashing the whole crawler.

Example code snippet for your parse method:

def parse_post(self, response):
    try:
        # Use default values to handle missing elements
        upvotes = response.css('.score::text').extract_first(default='0')
        title = response.css('h1._eYtD2XCVieq6emjKBH3m::text').extract_first(default='No Title')
        post_url = response.url
        
        yield {
            'title': title,
            'upvotes': upvotes,
            'url': post_url
        }
    except Exception as e:
        # Log the error instead of crashing
        self.logger.error(f"Failed to parse {response.url}: {str(e)}")
        # Optional: Yield an error entry to track problematic URLs
        yield {'error': str(e), 'url': response.url}

2. Reddit's Anti-Scraping Measures

Reddit actively blocks scrapers that don't mimic human browsing behavior. If you're hitting rate limits or sending requests without proper headers, the server might return malformed or empty responses, breaking your parser.

Fix: Tweak your settings.py to play nice with Reddit:

# Set a real user-agent to avoid being flagged as a bot
USER_AGENT = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'

# Add a delay between requests
DOWNLOAD_DELAY = 2

# Enable auto-throttling to adjust speed based on server response
AUTOTHROTTLE_ENABLED = True
AUTOTHROTTLE_START_DELAY = 1
AUTOTHROTTLE_MAX_DELAY = 5

3. Broken Pagination Logic

Your recursive pagination might be generating invalid after parameters, leading to 404 pages or malformed content that your crawler can't process.

Fix: Double-check how you're extracting the next page token. Here's a robust way to handle Reddit pagination:

def parse(self, response):
    # Parse all posts on the current page
    for post_container in response.css('div._1poyrkZ7g36PawDueRza-J'):
        post_link = post_container.css('a._3jOxDPIQ0KaOWpzvSQo-1s::attr(href)').extract_first()
        if post_link:
            yield response.follow(post_link, self.parse_post)
    
    # Get the next page URL safely
    next_page = response.css('span.next-button a::attr(href)').extract_first()
    if next_page:
        # Only follow if the URL is valid
        yield response.follow(next_page, self.parse)

4. Get the Full Error Traceback

Since your original error log was cut off, the exact issue might be something specific (like a KeyError from a missing dictionary key). To get the full picture:

Fix: Run your crawler with verbose logging to capture the complete traceback:

scrapy crawl your_spider_name --logfile=scrapy.log -L DEBUG

Open scrapy.log after the crash—this will show you exactly which line of code is failing and why.

内容的提问来源于stack exchange,提问作者jay3shh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:25:59