You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬虫仅爬至第3页停止,需解决50页分页爬取问题

Hey there! Let's troubleshoot why your Scrapy crawler stops at page 3 instead of getting all 50 pages. Based on your setup with Rule and LinkExtractor, here are the most likely issues and fixes:

It's common for sites to tweak URL structures or button attributes after a few pages, which breaks your link matching.

  • Check your URL pattern: If you used allow=r'/page/\d{1}', that only matches single-digit pages (1-9). Swap it for allow=r'/page/\d+' to catch all numeric page numbers.
  • Use precise element targeting: Instead of relying solely on URL patterns, target the "Next" button directly with XPath/CSS to avoid pattern mismatches. For example:
    Rule(LinkExtractor(restrict_xpaths='//a[normalize-space(text())="Next"]'), follow=True)
    
    The normalize-space() ensures extra spaces around "Next" don't break the match.

2. Anti-scraping measures are blocking you

Many sites throttle crawlers after a few requests, hiding the "Next" button or blocking access entirely. Try these fixes:

  • Add a download delay in settings.py:
    DOWNLOAD_DELAY = 2  # Wait 2 seconds between requests
    
  • Rotate user agents to avoid being flagged as a bot. You can use the scrapy-user-agents library to automate this.
  • Check if the site requires cookies or a session—some sites restrict pagination to users with active sessions, so enable cookies in your settings:
    COOKIES_ENABLED = True
    

3. Your Rule configuration has a subtle issue

Double-check your Rule parameters:

  • Ensure follow=True is set—this tells Scrapy to keep following links extracted by the Rule. If you have a callback, make sure it doesn't accidentally prevent follow behavior.
  • If your "Next" button loads content via JavaScript, Scrapy's default LinkExtractor won't see it. In this case, use tools like scrapy-playwright or scrapy-selenium to render JavaScript and extract the dynamic link.

4. Manual debugging steps to narrow it down

  • Use the Scrapy shell to inspect page 3 directly:
    scrapy shell https://your-site.com/page/3
    
    Run response.xpath('//a[contains(text(),"Next")]').get() to see if the link exists in the raw HTML. If it doesn't, the site is likely blocking you or using JS.
  • Enable debug logs in settings.py (LOG_LEVEL = 'DEBUG') to see if Scrapy is even trying to extract page 4's link. Look for lines starting with LinkExtractor to confirm links are being found or filtered.

Example fixed code

Here's a revised version of your spider with more robust pagination handling:

from scrapy.linkextractors import LinkExtractor
from scrapy.spiders import Rule, CrawlSpider

class MySpider(CrawlSpider):
    name = 'my_spider'
    start_urls = ['your_base_url_here']
    
    rules = (
        Rule(
            LinkExtractor(
                restrict_xpaths='//a[normalize-space(text())="Next"]',
                allow_domains=['your-site-domain.com']  # Prevent crawling external links
            ),
            callback='parse_item',
            follow=True
        ),
    )
    
    def parse_item(self, response):
        # Your existing item extraction logic goes here
        yield {
            'title': response.xpath('//h1/text()').get(),
            # Add other fields as needed
        }

内容的提问来源于stack exchange,提问作者Matthew Morrissey

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 12:28:16