You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy分页爬取异常:仅返回60条Startech显示器数据求助

问题排查与修复

核心问题:分页链接获取逻辑错误

你的爬虫仅返回60条结果,根源在于分页链接的XPath选择器无法正确定位「下一页」地址:
原代码中response.xpath("//ul[@class = 'pagination']/li/a/@href").get()只会提取匹配到的第一个链接,而这个链接通常是「首页」或「上一页」,并非真正的下一页入口,导致爬虫爬完前3页(每页20条)就停止遍历。

修复方案

根据目标网站的分页结构,下一页链接带有rel="next"属性,修改提取逻辑即可解决问题:

修复后的完整代码

import scrapy

class StartechSpider(scrapy.Spider):
    name = 'startech'
    allowed_domains = ['startech.com.bd']
    start_urls = ['https://www.startech.com.bd/monitor/']

    def parse(self, response):
        # 提取当前页显示器数据
        monitors = response.xpath("//div[@class='p-item']")
        for monitor in monitors:
            item = monitor.xpath(".//h4[@class='p-item-name']/a/text()").get()
            price = monitor.xpath(".//div[@class='p-item-price']/span/text()").get()
            
            yield {
                'item': item.strip() if item else None,
                'price': price.strip() if price else None
            }
            
        # 精准定位下一页链接:选择带rel="next"属性的a标签
        next_page = response.xpath("//ul[@class='pagination']/li/a[@rel='next']/@href").get()
        
        if next_page:
            yield response.follow(next_page, callback=self.parse)

额外优化建议

  • 数据清洗:对提取的商品名和价格做strip()处理,去除多余空格与换行
  • 反爬适配:在settings.py中设置DOWNLOAD_DELAY = 2,降低请求频率避免被限制
  • 进度跟踪:可添加打印语句输出当前爬取的页面URL,方便排查流程

内容的提问来源于stack exchange,提问作者Bashar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 20:50:33