Scrapy爬取时页面链接相同的情况下如何实现翻页?
Hey there! I’ve run into this exact issue before—when a site keeps the same URL but loads new content for subsequent pages (usually via AJAX, JS rendering, or hidden form parameters). Let’s break down how to solve this based on what’s happening under the hood of your target site.
First, Diagnose the Page Loading Mechanism
Before writing code, fire up your browser’s DevTools (Network tab) and click the "Next Page" button. Watch what requests are sent:
- Is there a POST request to the same URL with form data like
page=2oroffset=20? - Is there an XHR GET request to an API endpoint with pagination parameters?
- Or does the site just update the DOM via JavaScript without sending a new request?
Once you know that, pick the right approach below.
Solution 1: Pagination via Hidden Parameters (Most Common)
If the site uses form data or query parameters (even if the visible URL doesn’t change), you can manually construct pagination requests. Here’s how to adapt your existing parse method:
import scrapy class YourSpider(scrapy.Spider): name = "your_spider" start_urls = ["目标网站链接"] def parse(self, response): # Your existing scraping logic self.main_cat = response.xpath('//div[@id="products_content"]/div/text()').extract() self.sub_cat = response.xpath('//div[@class="accordion"]/div[@class="title"]/text()').extract() onclick_list = response.xpath('//div[@class="accordion"]/div[@class="no_title subtitle_chck"]/@onclick').extract() for index in range(len(onclick_list)): sub_sub_cat = response.xpath('...') # Your existing selector # Yield your item here, e.g.: # yield { # "main_cat": self.main_cat, # "sub_cat": self.sub_cat[index], # "sub_sub_cat": sub_sub_cat # } # Handle pagination current_page = response.meta.get("current_page", 1) next_page_num = current_page + 1 # Check if next page exists (adjust XPath to match your site's "Next" button/indicator) has_next_page = response.xpath('//button[contains(text(), "Next Page")]').exists() if has_next_page: # If using POST requests with form data (adjust form fields to match what you saw in DevTools) yield scrapy.FormRequest( url=response.url, formdata={"page": str(next_page_num), "limit": "10"}, # Example params meta={"current_page": next_page_num}, callback=self.parse, dont_filter=True # Critical: Disable default URL-based deduplication ) # If using GET requests with query params (even if URL looks same, backend may accept them) # yield scrapy.Request( # url=f"{response.url}?page={next_page_num}", # meta={"current_page": next_page_num}, # callback=self.parse, # dont_filter=True # )
Solution 2: JavaScript-Rendered Pages (No Explicit Requests)
If the site updates content purely via JS (no new network requests), you’ll need to simulate browser behavior to trigger the page load. Use scrapy-playwright or scrapy-splash for this. Here’s a Playwright example:
First, install scrapy-playwright:
pip install scrapy-playwright
Then update your spider:
from scrapy_playwright.page import PageCoroutine class YourSpider(scrapy.Spider): name = "your_spider" start_urls = ["目标网站链接"] def start_requests(self): for url in self.start_urls: yield scrapy.Request( url, meta=dict( playwright=True, playwright_page_coroutines=[ PageCoroutine("wait_for_selector", "div#products_content"), # Wait for content to load ], ), ) def parse(self, response): # Your existing scraping logic # ... (same as before) # Track items scraped to avoid infinite loops items_scraped = response.meta.get("items_scraped", 0) current_page_items = len(onclick_list) # Adjust to match your item count new_items_total = items_scraped + current_page_items # Check for next page button next_button_selector = '//button[contains(text(), "Next")]' has_next_page = response.xpath(next_button_selector).exists() # Stop if no new items are loaded (prevents infinite loops) if has_next_page and current_page_items > 0: yield scrapy.Request( response.url, meta=dict( playwright=True, playwright_page_coroutines=[ PageCoroutine("click", next_button_selector), # Click next page PageCoroutine("wait_for_selector", "div#products_content", state="visible"), # Wait for new content ], items_scraped=new_items_total ), callback=self.parse, dont_filter=True )
Key Notes to Avoid Headaches
- Always set
dont_filter=True: Scrapy’s default deduplication ignores requests with the same URL, so this disables that. - Add termination conditions: Check if new content is loaded (e.g., item count doesn’t increase) or if the "Next" button disappears—otherwise your spider will run forever.
- Respect robots.txt and rate limits: Dynamic pagination can trigger anti-bot measures, so add delays with
DOWNLOAD_DELAYin yoursettings.py.
内容的提问来源于stack exchange,提问作者Danyal Mughal

