pd.json_normalize函数record_path嵌套后meta嵌套数据失效问题求助
解决pd.json_normalize嵌套record_path下meta嵌套字段的问题
你的代码其实已经正确提取了meta中的嵌套字段,只是生成的列名是扁平化的点分隔格式(比如info.teachers.math),看起来像是没生效,实际功能是正常的。
你的代码运行结果
执行你提供的代码后,输出的DataFrame会包含info.teachers.math列,内容正确对应每个班级的数学老师:
name sex grades.math grades.physics class room info.teachers.math 0 Tom M 66.0 77.0 Year 1 Yellow Rick Scott 1 James M 80.0 78.0 Year 1 Yellow Rick Scott 2 Tony M NaN NaN Year 2 Blue Alan Turing 3 Jacqueline F NaN NaN Year 2 Blue Alan Turing
优化方案:让字段更清晰
如果觉得扁平化的列名不够直观,可以通过meta_prefix和record_prefix参数给元数据和学生数据的字段分别添加前缀,避免混淆:
import pandas as pd json_list = [ { 'class': 'Year 1', 'student count': 20, 'room': 'Yellow', 'info': { 'teachers': { 'math': 'Rick Scott', 'physics': 'Elon Mask' } }, 'extra': { 'students': [ { 'name': 'Tom', 'sex': 'M', 'grades': { 'math': 66, 'physics': 77 } }, { 'name': 'James', 'sex': 'M', 'grades': { 'math': 80, 'physics': 78 } }, ] } }, { 'class': 'Year 2', 'student count': 25, 'room': 'Blue', 'info': { 'teachers': { 'math': 'Alan Turing', 'physics': 'Albert Einstein' } }, 'extra': { 'students': [ { 'name': 'Tony', 'sex': 'M' }, { 'name': 'Jacqueline', 'sex': 'F' }, ] } }, ] df = pd.json_normalize( json_list, record_path=['extra', 'students'], meta=['class', 'room', ['info', 'teachers', 'math'], ['student count']], meta_prefix='class_', # 给班级相关元数据加前缀 record_prefix='student_' # 给学生数据加前缀 ) print(df)
输出结果列名更清晰:
student_name student_sex student_grades.math student_grades.physics class_class class_room class_info.teachers.math class_student count 0 Tom M 66.0 77.0 Year 1 Yellow Rick Scott 20 1 James M 80.0 78.0 Year 1 Yellow Rick Scott 20 2 Tony M NaN NaN Year 2 Blue Alan Turing 25 3 Jacqueline F NaN NaN Year 2 Blue Alan Turing 25
排查路径错误
如果确实出现meta字段未被提取的情况,检查路径拼写是否正确,或者添加errors='raise'参数触发异常,快速定位问题:
pd.json_normalize( json_list, record_path=['extra', 'students'], meta=['class', 'room', ['info', 'teachers', 'math']], errors='raise' )
内容的提问来源于stack exchange,提问作者DidSquids
相关产品推荐
相关产品推荐

