使用Scrapy爬取东京奥运运动员页面时如何用Xpath/CSS提取rank列数据
Scrapy无法提取rank字段问题解决
核心原因
- 你的XPath路径中包含
tbody节点:该节点是浏览器渲染页面时自动补全的,Scrapy获取的原始HTML源码不存在这个节点,因此路径匹配全程失效,返回None - 绝对层级的XPath稳定性极低,页面结构稍有调整就会匹配失败,且单独计数循环很容易出现event和rank顺序不匹配的问题
修正后的代码
把parse方法替换为以下内容即可:
def parse(self, response): # 先选中所有赛事对应的表格行 event_rows = response.xpath('//div[contains(@class,"table-responsive")]//table//tr[td/a[contains(@class,"eventTagLink")]]') event = [] rank = [] for row in event_rows: event.append(row.xpath('.//a[contains(@class,"eventTagLink")]/text()').get().strip()) # 提取排名,处理可能为空的情况 r = row.xpath('./td[3]/text()').get() rank.append(r.strip() if r else None) yield{ 'name' : response.css('h1::text').get().strip(), 'noc' : response.css('div.playerTag::attr(country)').get(), 'team' : response.css('a.country::text').get(), 'sport' : response.css('div.container-fluid > div.row > a::text').get(), 'sex' : response.xpath('//div/div[1]/div[1]/div[2]/div[1]/div/div[2]/div/div[3]/div[1]/div[3]/text()').extract()[-1].strip(), 'age': response.xpath('//div/div[1]/div[1]/div[2]/div[1]/div/div[2]/div/div[3]/div[1]/div[2]/text()').extract()[-1].strip(), 'event':event, 'rank':rank }
验证效果
目标页面测试返回的rank值为['25', '44'],和页面展示的运动员两项赛事排名完全一致。
内容的提问来源于stack exchange,提问作者David Molina
相关产品推荐
相关产品推荐

