如何在Scrapy中沿用Requests+BeautifulSoup的分页处理逻辑?
哈哈,完全理解你不想切换熟悉逻辑的心情!其实在Scrapy里完全能复刻你之前用Requests+BeautifulSoup那套分页思路,不用非得硬套response.follow()。下面给你一步步拆解怎么实现:
1. 用start_requests()初始化分页起始值
Scrapy爬虫类里可以重写start_requests()方法,在这里设置和你之前一致的初始页码,然后构造第一个请求,把当前页码存入meta参数,方便后续递增使用。
import scrapy from bs4 import BeautifulSoup # 保留你熟悉的BeautifulSoup解析方式 class MyPageSpider(scrapy.Spider): name = 'page_spider' # 替换成你实际的URL模板,和之前Requests里的格式保持一致 base_url = 'http://example.com/list?page={}' def start_requests(self): initial_page = 0 # 和你之前的初始页码完全一致 yield scrapy.Request( url=self.base_url.format(initial_page), callback=self.parse_page, meta={'current_page': initial_page} # 把当前页码传递给解析函数 )
2. 复刻BeautifulSoup解析逻辑
在parse_page解析函数里,你可以直接用response.text拿到页面内容,然后用BeautifulSoup做数据提取,这部分代码和你之前写的几乎完全一样。处理完当前页数据后,再按照你熟悉的逻辑递增页码、判断是否继续爬取下一页。
def parse_page(self, response): # 完全照搬你之前的BeautifulSoup解析逻辑 soup = BeautifulSoup(response.text, 'html.parser') items = soup.find_all('div', class_='target-item') # 替换成你实际的元素选择器 # 提取当前页数据,和之前的逻辑一致 for item in items: yield { 'title': item.find('h2').text.strip(), 'content': item.find('p').text.strip(), # 其他需要提取的字段... } # 分页逻辑:递增页码,判断是否继续请求 current_page = response.meta['current_page'] next_page = current_page + 1 # 这里替换成你之前的停止判断条件,比如: # - 判断当前页是否有数据(items不为空) # - 判断页面是否存在"下一页"按钮 # - 设置最大爬取页码上限 if len(items) > 0: yield scrapy.Request( url=self.base_url.format(next_page), callback=self.parse_page, meta={'current_page': next_page} )
3. 可选:用Scrapy自带选择器替代BeautifulSoup
如果你想更贴合Scrapy的生态,也可以把BeautifulSoup换成Scrapy自带的CSS/XPath选择器,但这一步完全看你的习惯——不想换就继续用BeautifulSoup也完全没问题。示例如下:
def parse_page(self, response): # 用Scrapy CSS选择器替代BeautifulSoup items = response.css('div.target-item') for item in items: yield { 'title': item.css('h2::text').get().strip(), 'content': item.css('p::text').get().strip(), } # 分页逻辑和上面完全一致 current_page = response.meta['current_page'] next_page = current_page + 1 if len(items) > 0: yield scrapy.Request( url=self.base_url.format(next_page), callback=self.parse_page, meta={'current_page': next_page} )
内容的提问来源于stack exchange,提问作者SIM
相关产品推荐
相关产品推荐

