使用Scrapy爬取英超球员页面异常,请求排查代码问题
以下是你的代码存在的核心问题及对应修复方案:
缩进错误,方法未绑定到类
你的parse和parse_players函数没有缩进在PlStatsSpider类内部,Scrapy无法识别这些作为爬虫的回调方法,会导致爬虫启动后无解析逻辑执行。
修复:将两个函数缩进,与类内的name、allowed_domains等属性对齐。起始URL错误,未指向球员列表页
你设置的start_urls是主页http://premierleague.com,但主页中并没有你要提取的//td/a链接(这些链接在/players页面),爬虫从主页开始后根本找不到任何球员链接。
修复:将start_urls改为球员列表页:['https://www.premierleague.com/players']URL拼接冗余且协议不一致
你手动用urljoin拼接URL,但response.follow可以直接接收相对链接,会自动基于当前响应的URL生成绝对地址;同时你用http协议拼接,目标网站实际是https,可能引发重定向或访问问题。
修复:直接传入相对链接给response.follow,去掉手动拼接逻辑:yield response.follow(link, callback=self.parse_players)XPath索引依赖易出错,未做空值处理
你的XPath使用固定索引(如[3]、[5]),一旦页面结构微调就会提取错误;且直接调用.strip()如果结果为None会抛出AttributeError。
修复:改用更精准的XPath定位(避免依赖索引),并添加空值判断后再处理。可能存在反爬拦截
英超官网可能拦截Scrapy默认的User-Agent,导致请求被拒绝或返回空页面。
修复:在settings.py中设置自定义User-Agent:USER_AGENT = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
修正后的完整代码
import scrapy class PlStatsSpider(scrapy.Spider): name = 'pl_stats' allowed_domains = ['premierleague.com'] start_urls = ['https://www.premierleague.com/players'] def parse(self, response): # 提取球员相对链接,直接传给response.follow自动生成绝对URL for link in response.xpath('//td/a/@href').getall(): yield response.follow(link, callback=self.parse_players) def parse_players(self, response): # 基于标签文本定位,避免依赖索引 dob_text = response.xpath('//div[@class="personalLists"]//div[contains(./div/text(), "Date of Birth")]/following-sibling::div/text()').get() height_text = response.xpath('//div[@class="personalLists"]//div[contains(./div/text(), "Height")]/following-sibling::div/text()').get() weight_text = response.xpath('//div[@class="personalLists"]//div[contains(./div/text(), "Weight")]/following-sibling::div/text()').get() position_text = response.xpath('//section[@class="sideWidget playerIntro t2-topBorder"]//div[contains(./div/text(), "Position")]/following-sibling::div/text()').get() club_text = response.xpath('//div[@class="info"]/a/text()').get() yield { 'Name': response.xpath('//h1/div[@class="name t-colour"]/text()').get(), 'DOB': dob_text.strip() if dob_text else None, 'Height': height_text, 'Club': club_text.strip() if club_text else None, 'Weight': weight_text, 'Position': position_text, 'Nationality': response.xpath('//span[@class="playerCountry"]/text()').get() }
内容的提问来源于stack exchange,提问作者hareko

