You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Scrapy爬取英超球员页面异常,请求排查代码问题

爬虫问题排查与修复

以下是你的代码存在的核心问题及对应修复方案:

  • 缩进错误,方法未绑定到类
    你的parse和parse_players函数没有缩进在PlStatsSpider类内部,Scrapy无法识别这些作为爬虫的回调方法,会导致爬虫启动后无解析逻辑执行。
    修复:将两个函数缩进,与类内的name、allowed_domains等属性对齐。

  • 起始URL错误,未指向球员列表页
    你设置的start_urls是主页http://premierleague.com,但主页中并没有你要提取的//td/a链接(这些链接在/players页面),爬虫从主页开始后根本找不到任何球员链接。
    修复:将start_urls改为球员列表页:['https://www.premierleague.com/players']

  • URL拼接冗余且协议不一致
    你手动用urljoin拼接URL,但response.follow可以直接接收相对链接,会自动基于当前响应的URL生成绝对地址;同时你用http协议拼接,目标网站实际是https,可能引发重定向或访问问题。
    修复:直接传入相对链接给response.follow,去掉手动拼接逻辑:yield response.follow(link, callback=self.parse_players)

  • XPath索引依赖易出错,未做空值处理
    你的XPath使用固定索引(如[3]、[5]),一旦页面结构微调就会提取错误;且直接调用.strip()如果结果为None会抛出AttributeError。
    修复:改用更精准的XPath定位(避免依赖索引),并添加空值判断后再处理。

  • 可能存在反爬拦截
    英超官网可能拦截Scrapy默认的User-Agent,导致请求被拒绝或返回空页面。
    修复:在settings.py中设置自定义User-Agent:

    USER_AGENT = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
    

修正后的完整代码

import scrapy


class PlStatsSpider(scrapy.Spider):
    name = 'pl_stats'
    allowed_domains = ['premierleague.com']
    start_urls = ['https://www.premierleague.com/players']

    def parse(self, response):
        # 提取球员相对链接,直接传给response.follow自动生成绝对URL
        for link in response.xpath('//td/a/@href').getall():
            yield response.follow(link, callback=self.parse_players)

    def parse_players(self, response):
        # 基于标签文本定位,避免依赖索引
        dob_text = response.xpath('//div[@class="personalLists"]//div[contains(./div/text(), "Date of Birth")]/following-sibling::div/text()').get()
        height_text = response.xpath('//div[@class="personalLists"]//div[contains(./div/text(), "Height")]/following-sibling::div/text()').get()
        weight_text = response.xpath('//div[@class="personalLists"]//div[contains(./div/text(), "Weight")]/following-sibling::div/text()').get()
        position_text = response.xpath('//section[@class="sideWidget playerIntro t2-topBorder"]//div[contains(./div/text(), "Position")]/following-sibling::div/text()').get()
        club_text = response.xpath('//div[@class="info"]/a/text()').get()

        yield {
            'Name': response.xpath('//h1/div[@class="name t-colour"]/text()').get(),
            'DOB': dob_text.strip() if dob_text else None,
            'Height': height_text,
            'Club': club_text.strip() if club_text else None,
            'Weight': weight_text,
            'Position': position_text,
            'Nationality': response.xpath('//span[@class="playerCountry"]/text()').get()
        }

内容的提问来源于stack exchange,提问作者hareko

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 22:05:29