You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Scrapy爬取sofifa.com异常:爬取超量且无法自动停止

问题分析与修复方案

核心问题原因

  1. follow=True导致爬虫无限爬取:原代码中Rule的follow=True会让Scrapy自动跟进每个响应里所有符合规则的链接,包括球员页面内的球队、其他球员等无关链接,导致爬虫无法停止且爬取大量额外内容。
  2. restrict_xpaths切片写法无效:('//table//tbody//tr/td[2]/a')[:60]是对字符串做切片,而非提取前60个元素,LinkExtractor会忽略这个操作,实际还是会提取页面内所有符合该XPath的链接。
  3. 链接过滤不彻底:虽然parse_item里判断了/player路径,但LinkExtractor仍会提取非球员链接并发送请求,造成资源浪费。

修复后的代码

import scrapy

class PlayersSpider(scrapy.Spider):
    name = "players"
    allowed_domains = ["sofifa.com"]
    user_agent = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/121.0.0.0 Safari/537.36'

    def start_requests(self):
        yield scrapy.Request(
            url='https://sofifa.com',
            headers={'User-Agent': self.user_agent},
            callback=self.parse_homepage
        )

    def parse_homepage(self, response):
        # 提取首页前60个球员的链接
        player_links = response.xpath('//table//tbody//tr/td[2]/a/@href')[:60]
        for link in player_links:
            player_url = response.urljoin(link.get())
            yield scrapy.Request(
                url=player_url,
                headers={'User-Agent': self.user_agent},
                callback=self.parse_player
            )

    def parse_player(self, response):
        yield {
            'full_name': response.xpath('//div[@class="profile clearfix"]/h1/text()').get().strip(),
            'overall_rating': response.xpath('//div[@class="grid"]//em[1]/text()').get().strip()
        }

修改说明

  • 改用普通scrapy.Spider替代CrawlSpider,完全掌控爬取流程,避免自动跟进无关链接。
  • 新增parse_homepage方法专门处理首页,直接对提取的链接列表做[:60]切片,确保只处理目标球员。
  • 移除follow=True,仅爬取首页指定的60个球员链接,爬虫完成后自动停止。
  • 对提取的文本做.strip()处理,清除多余空格,优化数据格式。
  • 若需防反爬,建议在settings.py中设置DOWNLOAD_DELAY = 1替代硬编码的time.sleep。

内容的提问来源于stack exchange,提问作者Muaaz Abu Zaid

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 00:57:07