You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Extruct提取JSON-LD格式SportsEvent节点值为空问题排查

问题根因

extruct提取的json-ld结构中,schema类型通过@type字段标识,你的原有判断是检查字典中是否存在SportsEvent键,和实际结构不匹配。目标站点的SportsEvent节点结构如下:

{
  "@context": "https://schema.org",
  "@type": "SportsEvent",
  "awayTeam": {"@type":"SportsTeam", "name":"xxx"},
  "homeTeam": {"@type":"SportsTeam", "name":"xxx"},
  "location": {"@type":"Place", "name":"xxx"},
  "startDate": "xxxx-xx-xxTxx:xx:xx"
}
修复后的核心代码

只需要替换save_teams函数即可正常提取数据:

def save_teams(data, df):
    for item in data['json-ld']:
        # 兼容@type为字符串或列表的情况
        item_type = item.get('@type', '')
        is_sports_event = (item_type == 'SportsEvent') or (isinstance(item_type, list) and 'SportsEvent' in item_type)
        if not is_sports_event:
            continue
        # 符合条件的item本身就是SportsEvent对象,直接取字段
        row = {
            'awayTeam': item.get('awayTeam', {}).get('name'),
            'homeTeam': item.get('homeTeam', {}).get('name'),
            'location': item.get('location', {}).get('name'),
            'startDate': item.get('startDate')
        }
        # 替换废弃的append方法,兼容新旧pandas版本
        df = pd.concat([df, pd.DataFrame([row])], ignore_index=True)
    return df
额外优化建议
  • 你可以直接移除get_next_page相关的分页逻辑,目标站点的odds页面没有分页,这部分代码不会生效只会重复发起请求
  • 可以复用同一个webdriver实例,不需要每次请求都重新创建,能大幅提升爬取速度

内容的提问来源于stack exchange,提问作者Leo Torres

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 23:06:07