You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Scrapy爬取东京奥运运动员页面时如何用Xpath/CSS提取rank列数据

Scrapy无法提取rank字段问题解决

核心原因

  • 你的XPath路径中包含tbody节点:该节点是浏览器渲染页面时自动补全的,Scrapy获取的原始HTML源码不存在这个节点,因此路径匹配全程失效,返回None
  • 绝对层级的XPath稳定性极低,页面结构稍有调整就会匹配失败,且单独计数循环很容易出现event和rank顺序不匹配的问题

修正后的代码

把parse方法替换为以下内容即可:

def parse(self, response):
    # 先选中所有赛事对应的表格行
    event_rows = response.xpath('//div[contains(@class,"table-responsive")]//table//tr[td/a[contains(@class,"eventTagLink")]]')
    event = []
    rank = []
    for row in event_rows:
        event.append(row.xpath('.//a[contains(@class,"eventTagLink")]/text()').get().strip())
        # 提取排名,处理可能为空的情况
        r = row.xpath('./td[3]/text()').get()
        rank.append(r.strip() if r else None)
                    
    yield{
        'name' : response.css('h1::text').get().strip(),
        'noc' : response.css('div.playerTag::attr(country)').get(),
        'team' : response.css('a.country::text').get(),
        'sport' : response.css('div.container-fluid > div.row > a::text').get(),
        'sex' : response.xpath('//div/div[1]/div[1]/div[2]/div[1]/div/div[2]/div/div[3]/div[1]/div[3]/text()').extract()[-1].strip(),
        'age': response.xpath('//div/div[1]/div[1]/div[2]/div[1]/div/div[2]/div/div[3]/div[1]/div[2]/text()').extract()[-1].strip(), 
        'event':event,
        'rank':rank
    }

验证效果

目标页面测试返回的rank值为['25', '44'],和页面展示的运动员两项赛事排名完全一致。

内容的提问来源于stack exchange,提问作者David Molina

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 23:21:01