使用Extruct提取JSON-LD格式SportsEvent节点值为空问题排查
问题根因
extruct提取的json-ld结构中,schema类型通过@type字段标识,你的原有判断是检查字典中是否存在SportsEvent键,和实际结构不匹配。目标站点的SportsEvent节点结构如下:
{ "@context": "https://schema.org", "@type": "SportsEvent", "awayTeam": {"@type":"SportsTeam", "name":"xxx"}, "homeTeam": {"@type":"SportsTeam", "name":"xxx"}, "location": {"@type":"Place", "name":"xxx"}, "startDate": "xxxx-xx-xxTxx:xx:xx" }
修复后的核心代码
只需要替换save_teams函数即可正常提取数据:
def save_teams(data, df): for item in data['json-ld']: # 兼容@type为字符串或列表的情况 item_type = item.get('@type', '') is_sports_event = (item_type == 'SportsEvent') or (isinstance(item_type, list) and 'SportsEvent' in item_type) if not is_sports_event: continue # 符合条件的item本身就是SportsEvent对象,直接取字段 row = { 'awayTeam': item.get('awayTeam', {}).get('name'), 'homeTeam': item.get('homeTeam', {}).get('name'), 'location': item.get('location', {}).get('name'), 'startDate': item.get('startDate') } # 替换废弃的append方法,兼容新旧pandas版本 df = pd.concat([df, pd.DataFrame([row])], ignore_index=True) return df
额外优化建议
- 你可以直接移除
get_next_page相关的分页逻辑,目标站点的odds页面没有分页,这部分代码不会生效只会重复发起请求 - 可以复用同一个webdriver实例,不需要每次请求都重新创建,能大幅提升爬取速度
内容的提问来源于stack exchange,提问作者Leo Torres
相关产品推荐
相关产品推荐

