You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy分页异常:仅能爬取前2页,无法获取剩余页面数据

问题

爬取cdw.com网站搜索"axiom"的结果,目标共17页,但Scrapy脚本仅能获取第1、2页数据,爬虫返回前两页后自行关闭,需解决获取剩余15页数据的问题。

原代码片段
import scrapy  
from cdwfinal.items import CdwfinalItem  
from scrapy.selector import Selector  
import datetime  
import pandas as pd  
import time  


class CdwSpider(scrapy.Spider):  
    name = 'cdw'
    allowed_domains = ['www.cdw.com']
    start_urls = ['http://www.cdw.com/']
    base_url = 'http://www.cdw.com'


    def start_requests(self):
   
        yield scrapy.Request(url = 'https://www.cdw.com/search/?key=axiom' , callback=self.parse )
    
    def parse(self, response): 
    
        item=[]
        hxs = Selector(response)
        item = CdwfinalItem()
    
        abc = hxs.xpath('//*[@id="main"]//*[@class="grid-row"]')
    
        for i in range(len(abc)):
        
            try:
                item['mpn'] = hxs.xpath("//div[contains(@class,'search-results')]/div[contains(@class,'search-result')]["+ str(i+1) +"]//*[@class='mfg-code']/text()").extract()
            except:
                item['mpn'] = 'NA'

            try:
                item['part_no'] = hxs.xpath("//div[contains(@class,'search-results')]/div[contains(@class,'search-result')]["+ str(i+1) +"]//*[@class='cdw-code']/text()").extract()
            except:
                item['part_no'] = 'NA'

        
            
            yield item
    
        next_page = hxs.xpath('//*[@id="main"]//*[@class="no-hover" and @aria-label="Next Page"]').extract()
        if next_page:
            new_page_href =  hxs.xpath('//*[@id="main"]//*[@class="no-hover" and @aria-label="Next Page"]/@href').extract_first()
            new_page_url = response.urljoin(new_page_href)
            yield scrapy.Request(new_page_url, callback=self.parse, meta={"searchword": '123'})
运行日志
2023-02-11 15:39:55 [scrapy_user_agents.middlewares] DEBUG: Assigned User-Agent Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/53.0.2785.116 Safari/537.36
2023-02-11 15:39:55 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://www.cdw.com/search/?key=axiom&pcurrent=3> (referer: https://www.cdw.com/search/?key=axiom&pcurrent=2) ['cached']
2023-02-11 15:39:55 [scrapy.core.engine] INFO: Closing spider (finished)
2023-02-11 15:39:55 [scrapy.extensions.feedexport] INFO: Stored csv feed (48 items) in: Test5.csv
2023-02-11 15:39:55 [scrapy.statscollectors] INFO: Dumping Scrapy stats:
{'downloader/request_bytes': 2178,
 'downloader/request_count': 3,
 'downloader/request_method_count/GET': 3,
 'downloader/response_bytes': 68059,
 'downloader/response_count': 3,
 'downloader/response_status_count/200': 3,
 'elapsed_time_seconds': 1.30903,
 'feedexport/success_count/FileFeedStorage': 1,
 'finish_reason': 'finished',
 'finish_time': datetime.datetime(2023, 2, 11, 10, 9, 55, 327740),
 'httpcache/hit': 3,
 'httpcompression/response_bytes': 384267,
 'httpcompression/response_count': 3,
 'item_scraped_count': 48,
 'log_count/DEBUG': 62,
 'log_count/INFO': 11,
 'log_count/WARNING': 45,
 'request_depth_max': 2,
 'response_received_count': 3,
 'scheduler/dequeued': 3,
 'scheduler/dequeued/memory': 3,
 'scheduler/enqueued': 3,
 'scheduler/enqueued/memory': 3,
 'start_time': datetime.datetime(2023, 2, 11, 10, 9, 54, 18710)}
解决方案

核心问题分析

从日志可见,爬虫请求了第3页但标记为cached,且未继续请求后续页面。根源是下一页的XPath选择器依赖不稳定的class属性,当页面切换后,"Next Page"按钮的class可能变化(比如no-hover只在特定分页状态下存在),导致无法定位到后续分页链接;同时Item对象初始化逻辑有误,会导致数据覆盖。

1. 修正下一页XPath定位

去掉对no-hoverclass的依赖,改用更稳定的aria-label属性定位:

# 修改后的下一页定位逻辑
next_page = hxs.xpath('//*[@id="main"]//*[@aria-label="Next Page"]')
if next_page:
    new_page_href = hxs.xpath('//*[@id="main"]//*[@aria-label="Next Page"]/@href').extract_first()
    new_page_url = response.urljoin(new_page_href)
    yield scrapy.Request(new_page_url, callback=self.parse)

2. 修复Item重复赋值问题

将CdwfinalItem()初始化移到循环内部,避免所有Item共享同一个实例导致数据覆盖;同时用extract_first()代替extract(),直接获取单个文本值,简化异常处理:

# 修改后的Item处理逻辑
for i in range(len(abc)):
    item = CdwfinalItem()  # 移到循环内,每个条目生成独立实例
    # 获取mpn,找不到则返回NA
    item['mpn'] = hxs.xpath("//div[contains(@class,'search-results')]/div[contains(@class,'search-result')]["+ str(i+1) +"]//*[@class='mfg-code']/text()").extract_first() or 'NA'
    # 获取part_no,找不到则返回NA
    item['part_no'] = hxs.xpath("//div[contains(@class,'search-results')]/div[contains(@class,'search-result')]["+ str(i+1) +"]//*[@class='cdw-code']/text()").extract_first() or 'NA'
    yield item

3. 关闭缓存(临时排查)

日志中的cached说明Scrapy使用了缓存,可能导致获取的页面不是最新状态。在settings.py中关闭缓存:

HTTPCACHE_ENABLED = False

或运行爬虫时添加参数:

scrapy crawl cdw --no-cache

4. 添加下载延迟避免反爬

cdw.com可能存在反爬机制,过快请求会被拦截。在settings.py中添加:

DOWNLOAD_DELAY = 2  # 每2秒请求一次
修改后的完整代码
import scrapy  
from cdwfinal.items import CdwfinalItem  
from scrapy.selector import Selector  
import datetime  
import pandas as pd  
import time  


class CdwSpider(scrapy.Spider):  
    name = 'cdw'
    allowed_domains = ['www.cdw.com']
    start_urls = ['http://www.cdw.com/']
    base_url = 'http://www.cdw.com'


    def start_requests(self):
        yield scrapy.Request(url='https://www.cdw.com/search/?key=axiom', callback=self.parse )
    
    def parse(self, response): 
        hxs = Selector(response)
        abc = hxs.xpath('//*[@id="main"]//*[@class="grid-row"]')
    
        for i in range(len(abc)):
            item = CdwfinalItem()
            # 获取mpn
            item['mpn'] = hxs.xpath("//div[contains(@class,'search-results')]/div[contains(@class,'search-result')]["+ str(i+1) +"]//*[@class='mfg-code']/text()").extract_first() or 'NA'
            # 获取part_no
            item['part_no'] = hxs.xpath("//div[contains(@class,'search-results')]/div[contains(@class,'search-result')]["+ str(i+1) +"]//*[@class='cdw-code']/text()").extract_first() or 'NA'
            
            yield item
    
        # 定位下一页
        next_page = hxs.xpath('//*[@id="main"]//*[@aria-label="Next Page"]')
        if next_page:
            new_page_href = hxs.xpath('//*[@id="main"]//*[@aria-label="Next Page"]/@href').extract_first()
            new_page_url = response.urljoin(new_page_href)
            yield scrapy.Request(new_page_url, callback=self.parse)

内容的提问来源于stack exchange,提问作者Vinay Singh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 11:35:23