You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy下一页数据获取与页面跳转及多页爬取无输出排查

Hey there! Let's tackle your two Scrapy questions one by one, with practical fixes and explanations:

问题1:如何在Scrapy中获取下一页数据并跳转至其他页面?

Handling pagination in Scrapy is straightforward, and there are two common approaches depending on your needs:

方法1:手动提取下一页链接(灵活自定义)

This is the most flexible option, great for when you need custom logic around pagination:

  • First, process the data from the current page in your parse (or custom callback) method
  • Use XPath/CSS selectors to extract the next page's URL. For example, if the next page button has a class like next-page:
next_page_link = response.css('a.next-page::attr(href)').get()
  • Convert relative URLs to absolute ones with response.urljoin() (critical if the link doesn't include the full domain)
  • Send a new request to the next page, reusing your parse method to loop through pages:
if next_page_link:
    yield Request(url=response.urljoin(next_page_link), callback=self.parse)

方法2:用CrawlSpider自动跟进链接(简化固定规则)

If your pagination follows a consistent pattern, CrawlSpider can automate link following:

  • Import the necessary classes and define a Rule to match pagination links:
from scrapy.spiders import CrawlSpider, Rule
from scrapy.linkextractors import LinkExtractor

class SaabCrawlSpider(CrawlSpider):
    name = 'saab_crawler'
    allowed_domains = ['thesaabsite.com']
    start_urls = ['http://www.thesaabsite.com/parts_om.php']

    rules = (
        # Follow all pagination links and process each page with parse_item
        Rule(LinkExtractor(allow=r'/parts_om\.php\?page=\d+'), callback='parse_item', follow=True),
    )

    def parse_item(self, response):
        # Process data from each page here
        print("Processing page:", response.url)

问题2:排查Spider无输出的问题

Looking at your provided code, there are several clear issues causing the lack of output. Let's fix them step by step:

Issue 1: Invalid allowed_domains configuration

allowed_domains should only contain the base domain, not a full path. Your current value ['thesaabsite.com/parts_om.php'] will block all requests because the domain doesn't match. Correct it to:

allowed_domains = ['thesaabsite.com']

Issue 2: Truncated start_urls

Your start URL is incomplete (http://www.thesaabsite.com/parts_...), so Scrapy can't send a valid request. Replace it with the full, working URL (e.g., http://www.thesaabsite.com/parts_om.php).

Issue 3: Missing parse method

You didn't define a parse callback. Scrapy's default parse method does nothing, so even if requests succeed, there's no code to generate output. Add a parse method to handle responses and print content.

Fixed Full Code Example

# -*- coding: utf-8 -*-
from scrapy import Spider
from scrapy.http import Request

class SaabSpider(Spider):
    name = 'saab'
    allowed_domains = ['thesaabsite.com']
    # Use the complete, valid start URL
    start_urls = ['http://www.thesaabsite.com/parts_om.php']

    def parse(self, response):
        # Add your print logic here to confirm the spider is working
        print("Successfully fetched page:", response.url)
        print("Page title:", response.css('title::text').get())

        # Optional: Add pagination logic here (using the method from Question 1)
        # next_page = response.css('a.next-page::attr(href)').get()
        # if next_page:
        #     yield Request(url=response.urljoin(next_page), callback=self.parse)

Extra Checks if Still No Output

  • If the site blocks Scrapy via robots.txt, set ROBOTSTXT_OBEY = False in your settings.py to test
  • Verify you can access the target URL from your network (or set up a proxy if needed)
  • Double-check that your selectors match the actual page structure if you're extracting data later

内容的提问来源于stack exchange,提问作者Muhammad Danial

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:30:55