You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:Scrapy爬虫分页失效问题及爬虫优化需求

Fixing Scrapy Pagination in Your parse_items Method

Hey there! Let's get that pagination working properly for your Scrapy spider. Since you're already successfully scraping the first page's 50 items in parse_items, we can extend this method to handle subsequent pages seamlessly—no need to split your logic into separate functions.

Here are the most common solutions based on how the website implements pagination:

1. Pagination via URL Parameters (e.g., ?page=2)

Most sites use a simple page number parameter in the URL. Here's how to loop through pages until there's no more data:

import scrapy

class YourSpider(scrapy.Spider):
    name = "your_spider"
    start_urls = ["https://example.com/items"]  # Replace with your target first page URL

    def parse_items(self, response):
        # Step 1: Scrape all 50 item links from the current page
        item_links = response.css("your-item-link-selector::attr(href)").getall()
        for link in item_links:
            # Yield a request to scrape each item's details (adjust callback as needed)
            yield scrapy.Request(
                url=response.urljoin(link),
                callback=self.parse_item_detail
            )

        # Step 2: Handle pagination
        # Get current page number from URL (adjust logic if your URL format differs)
        if "page=" in response.url:
            current_page = int(response.url.split("page=")[-1])
        else:
            current_page = 1  # First page has no page parameter

        next_page = current_page + 1

        # Only proceed if the current page had the full 50 items (indicates more pages exist)
        if len(item_links) == 50:
            # Construct next page URL (modify base URL if needed)
            base_url = response.url.split("?")[0]
            next_page_url = f"{base_url}?page={next_page}"
            # Yield request for next page, using the same parse_items callback
            yield scrapy.Request(
                url=next_page_url,
                callback=self.parse_items
            )

    def parse_item_detail(self, response):
        # Your logic to scrape item details goes here
        yield {
            "title": response.css("h1::text").get(),
            # Add other fields you need to scrape...
        }

If the site uses a visible "Next" button instead of URL parameters, extract the button's href and follow it until it no longer exists:

def parse_items(self, response):
    # Step 1: Scrape current page items (same as above)
    item_links = response.css("your-item-link-selector::attr(href)").getall()
    for link in item_links:
        yield scrapy.Request(response.urljoin(link), callback=self.parse_item_detail)

    # Step 2: Extract next page link from the "Next" button
    next_page_link = response.css("a.next-page-button::attr(href)").get()
    # Check if next page link exists (stop if it doesn't)
    if next_page_link:
        yield scrapy.Request(
            url=response.urljoin(next_page_link),
            callback=self.parse_items
        )

Key Tips to Avoid Issues

  • Anti-scraping safeguards: Some sites block repeated requests. Add a realistic USER_AGENT in settings.py and set DOWNLOAD_DELAY = 2 to mimic human browsing.
  • Edge case handling: If the last page has fewer than 50 items, the first method will stop automatically (since len(item_links) won't equal 50).
  • Manual testing first: Visit the site's second/third pages to confirm the URL pattern or next button selector before coding—this saves you time debugging.

内容的提问来源于stack exchange,提问作者Samsul Islam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 07:11:06