如何将多层嵌套赛事JSON全量解析为pandas DataFrame
实现方法
你当前的代码硬编码了record_path=['games', 'tour_1'],因此只能固定读取第一轮的赛事数据。games字段本身是键为轮次名、值为对应场次列表的字典,不需要提前知道有多少轮、每轮叫什么名字,直接遍历字典即可读取全量轮次数据,以下是可直接复用的实现:
单文件全轮次读取
两种实现方式可选,第一种性能更好、逻辑更直观:
- 方式1:预处理比赛数据后直接构造DataFrame
import json import pandas as pd with open('tournament_7.json', encoding='utf-8') as data_file: data = json.load(data_file) all_matches = [] # 遍历所有轮次,不需要硬编码轮次名称 for tour_name, match_list in data['games'].items(): for match in match_list: # 给每场比赛补充所属轮次字段 match['tour_name'] = tour_name all_matches.append(match) # 生成比赛明细DataFrame df = pd.DataFrame(all_matches) # 追加锦标赛层级的元信息字段 meta_fields = ['name','start_date', 'end_date', 'tours', 'type', 'winner'] for field in meta_fields: df[field] = data[field]
- 方式2:沿用
json_normalize动态拼接各轮次数据
如果你希望保持原有json_normalize的调用习惯,可以遍历所有轮次分别读取后纵向拼接:
df_list = [] meta_fields = ['name','start_date', 'end_date', 'tours', 'type', 'winner'] for tour_name in data['games'].keys(): tour_df = pd.json_normalize( data, record_path=['games', tour_name], meta=meta_fields ) tour_df['tour_name'] = tour_name df_list.append(tour_df) df = pd.concat(df_list, ignore_index=True)
上千份文件批量聚合
批量处理时不要逐文件逐行往最终DataFrame追加数据,先把每个文件解析得到的DataFrame存入列表,最后一次性合并,性能会提升数十倍,参考代码如下:
from pathlib import Path # 替换为你的JSON文件存放目录路径 data_folder = Path('./your_data_dir') all_df_list = [] meta_fields = ['name','start_date', 'end_date', 'tours', 'type', 'winner'] # 自动匹配所有命名符合tournament_*.json规则的文件 for json_path in data_folder.glob('tournament_*.json'): with open(json_path, encoding='utf-8') as f: data = json.load(f) # 复用单文件解析逻辑 file_matches = [] for tour_name, match_list in data['games'].items(): for match in match_list: match['tour_name'] = tour_name file_matches.append(match) file_df = pd.DataFrame(file_matches) for field in meta_fields: file_df[field] = data[field] all_df_list.append(file_df) # 一次性合并所有文件的赛事数据 final_total_df = pd.concat(all_df_list, ignore_index=True)
提示:读取文件时建议显式指定
encoding='utf-8',避免跨系统默认编码差异导致的特殊字符解析报错。
内容的提问来源于stack exchange,提问作者Ismail
相关产品推荐
相关产品推荐

