Scrapy分页爬取异常:仅返回60条Startech显示器数据求助
问题排查与修复
核心问题:分页链接获取逻辑错误
你的爬虫仅返回60条结果,根源在于分页链接的XPath选择器无法正确定位「下一页」地址:
原代码中response.xpath("//ul[@class = 'pagination']/li/a/@href").get()只会提取匹配到的第一个链接,而这个链接通常是「首页」或「上一页」,并非真正的下一页入口,导致爬虫爬完前3页(每页20条)就停止遍历。
修复方案
根据目标网站的分页结构,下一页链接带有rel="next"属性,修改提取逻辑即可解决问题:
修复后的完整代码
import scrapy class StartechSpider(scrapy.Spider): name = 'startech' allowed_domains = ['startech.com.bd'] start_urls = ['https://www.startech.com.bd/monitor/'] def parse(self, response): # 提取当前页显示器数据 monitors = response.xpath("//div[@class='p-item']") for monitor in monitors: item = monitor.xpath(".//h4[@class='p-item-name']/a/text()").get() price = monitor.xpath(".//div[@class='p-item-price']/span/text()").get() yield { 'item': item.strip() if item else None, 'price': price.strip() if price else None } # 精准定位下一页链接:选择带rel="next"属性的a标签 next_page = response.xpath("//ul[@class='pagination']/li/a[@rel='next']/@href").get() if next_page: yield response.follow(next_page, callback=self.parse)
额外优化建议
- 数据清洗:对提取的商品名和价格做
strip()处理,去除多余空格与换行 - 反爬适配:在
settings.py中设置DOWNLOAD_DELAY = 2,降低请求频率避免被限制 - 进度跟踪:可添加打印语句输出当前爬取的页面URL,方便排查流程
内容的提问来源于stack exchange,提问作者Bashar
相关产品推荐
相关产品推荐

