You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Scrapy中沿用Requests+BeautifulSoup的分页处理逻辑?

哈哈,完全理解你不想切换熟悉逻辑的心情!其实在Scrapy里完全能复刻你之前用Requests+BeautifulSoup那套分页思路,不用非得硬套response.follow()。下面给你一步步拆解怎么实现:

1. 用start_requests()初始化分页起始值

Scrapy爬虫类里可以重写start_requests()方法,在这里设置和你之前一致的初始页码,然后构造第一个请求,把当前页码存入meta参数,方便后续递增使用。

import scrapy
from bs4 import BeautifulSoup  # 保留你熟悉的BeautifulSoup解析方式

class MyPageSpider(scrapy.Spider):
    name = 'page_spider'
    # 替换成你实际的URL模板,和之前Requests里的格式保持一致
    base_url = 'http://example.com/list?page={}'

    def start_requests(self):
        initial_page = 0  # 和你之前的初始页码完全一致
        yield scrapy.Request(
            url=self.base_url.format(initial_page),
            callback=self.parse_page,
            meta={'current_page': initial_page}  # 把当前页码传递给解析函数
        )

2. 复刻BeautifulSoup解析逻辑

在parse_page解析函数里,你可以直接用response.text拿到页面内容,然后用BeautifulSoup做数据提取,这部分代码和你之前写的几乎完全一样。处理完当前页数据后,再按照你熟悉的逻辑递增页码、判断是否继续爬取下一页。

def parse_page(self, response):
        # 完全照搬你之前的BeautifulSoup解析逻辑
        soup = BeautifulSoup(response.text, 'html.parser')
        items = soup.find_all('div', class_='target-item')  # 替换成你实际的元素选择器

        # 提取当前页数据,和之前的逻辑一致
        for item in items:
            yield {
                'title': item.find('h2').text.strip(),
                'content': item.find('p').text.strip(),
                # 其他需要提取的字段...
            }

        # 分页逻辑:递增页码,判断是否继续请求
        current_page = response.meta['current_page']
        next_page = current_page + 1

        # 这里替换成你之前的停止判断条件,比如:
        # - 判断当前页是否有数据(items不为空)
        # - 判断页面是否存在"下一页"按钮
        # - 设置最大爬取页码上限
        if len(items) > 0:
            yield scrapy.Request(
                url=self.base_url.format(next_page),
                callback=self.parse_page,
                meta={'current_page': next_page}
            )

3. 可选:用Scrapy自带选择器替代BeautifulSoup

如果你想更贴合Scrapy的生态,也可以把BeautifulSoup换成Scrapy自带的CSS/XPath选择器,但这一步完全看你的习惯——不想换就继续用BeautifulSoup也完全没问题。示例如下:

def parse_page(self, response):
        # 用Scrapy CSS选择器替代BeautifulSoup
        items = response.css('div.target-item')
        for item in items:
            yield {
                'title': item.css('h2::text').get().strip(),
                'content': item.css('p::text').get().strip(),
            }

        # 分页逻辑和上面完全一致
        current_page = response.meta['current_page']
        next_page = current_page + 1

        if len(items) > 0:
            yield scrapy.Request(
                url=self.base_url.format(next_page),
                callback=self.parse_page,
                meta={'current_page': next_page}
            )

内容的提问来源于stack exchange,提问作者SIM

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:26:48