You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬取时页面链接相同的情况下如何实现翻页?

Handling Pagination with Duplicate URLs in Scrapy

Hey there! I’ve run into this exact issue before—when a site keeps the same URL but loads new content for subsequent pages (usually via AJAX, JS rendering, or hidden form parameters). Let’s break down how to solve this based on what’s happening under the hood of your target site.

First, Diagnose the Page Loading Mechanism

Before writing code, fire up your browser’s DevTools (Network tab) and click the "Next Page" button. Watch what requests are sent:

  • Is there a POST request to the same URL with form data like page=2 or offset=20?
  • Is there an XHR GET request to an API endpoint with pagination parameters?
  • Or does the site just update the DOM via JavaScript without sending a new request?

Once you know that, pick the right approach below.

Solution 1: Pagination via Hidden Parameters (Most Common)

If the site uses form data or query parameters (even if the visible URL doesn’t change), you can manually construct pagination requests. Here’s how to adapt your existing parse method:

import scrapy

class YourSpider(scrapy.Spider):
    name = "your_spider"
    start_urls = ["目标网站链接"]

    def parse(self, response):
        # Your existing scraping logic
        self.main_cat = response.xpath('//div[@id="products_content"]/div/text()').extract()
        self.sub_cat = response.xpath('//div[@class="accordion"]/div[@class="title"]/text()').extract()
        onclick_list = response.xpath('//div[@class="accordion"]/div[@class="no_title subtitle_chck"]/@onclick').extract()
        
        for index in range(len(onclick_list)):
            sub_sub_cat = response.xpath('...')  # Your existing selector
            # Yield your item here, e.g.:
            # yield {
            #     "main_cat": self.main_cat,
            #     "sub_cat": self.sub_cat[index],
            #     "sub_sub_cat": sub_sub_cat
            # }

        # Handle pagination
        current_page = response.meta.get("current_page", 1)
        next_page_num = current_page + 1

        # Check if next page exists (adjust XPath to match your site's "Next" button/indicator)
        has_next_page = response.xpath('//button[contains(text(), "Next Page")]').exists()
        
        if has_next_page:
            # If using POST requests with form data (adjust form fields to match what you saw in DevTools)
            yield scrapy.FormRequest(
                url=response.url,
                formdata={"page": str(next_page_num), "limit": "10"},  # Example params
                meta={"current_page": next_page_num},
                callback=self.parse,
                dont_filter=True  # Critical: Disable default URL-based deduplication
            )

            # If using GET requests with query params (even if URL looks same, backend may accept them)
            # yield scrapy.Request(
            #     url=f"{response.url}?page={next_page_num}",
            #     meta={"current_page": next_page_num},
            #     callback=self.parse,
            #     dont_filter=True
            # )

Solution 2: JavaScript-Rendered Pages (No Explicit Requests)

If the site updates content purely via JS (no new network requests), you’ll need to simulate browser behavior to trigger the page load. Use scrapy-playwright or scrapy-splash for this. Here’s a Playwright example:

First, install scrapy-playwright:

pip install scrapy-playwright

Then update your spider:

from scrapy_playwright.page import PageCoroutine

class YourSpider(scrapy.Spider):
    name = "your_spider"
    start_urls = ["目标网站链接"]

    def start_requests(self):
        for url in self.start_urls:
            yield scrapy.Request(
                url,
                meta=dict(
                    playwright=True,
                    playwright_page_coroutines=[
                        PageCoroutine("wait_for_selector", "div#products_content"),  # Wait for content to load
                    ],
                ),
            )

    def parse(self, response):
        # Your existing scraping logic
        # ... (same as before)

        # Track items scraped to avoid infinite loops
        items_scraped = response.meta.get("items_scraped", 0)
        current_page_items = len(onclick_list)  # Adjust to match your item count
        new_items_total = items_scraped + current_page_items

        # Check for next page button
        next_button_selector = '//button[contains(text(), "Next")]'
        has_next_page = response.xpath(next_button_selector).exists()

        # Stop if no new items are loaded (prevents infinite loops)
        if has_next_page and current_page_items > 0:
            yield scrapy.Request(
                response.url,
                meta=dict(
                    playwright=True,
                    playwright_page_coroutines=[
                        PageCoroutine("click", next_button_selector),  # Click next page
                        PageCoroutine("wait_for_selector", "div#products_content", state="visible"),  # Wait for new content
                    ],
                    items_scraped=new_items_total
                ),
                callback=self.parse,
                dont_filter=True
            )

Key Notes to Avoid Headaches

  • Always set dont_filter=True: Scrapy’s default deduplication ignores requests with the same URL, so this disables that.
  • Add termination conditions: Check if new content is loaded (e.g., item count doesn’t increase) or if the "Next" button disappears—otherwise your spider will run forever.
  • Respect robots.txt and rate limits: Dynamic pagination can trigger anti-bot measures, so add delays with DOWNLOAD_DELAY in your settings.py.

内容的提问来源于stack exchange,提问作者Danyal Mughal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 08:31:23