将Pandas DataFrame类JSON字符串列转为作者信息DataFrame
解决方案
步骤1:将字符串格式的类JSON数据转为Python对象
你的authorships列是字符串类型的数组,首先需要把每个单元格的内容转换成可操作的Python列表(内含字典)。用json.loads配合pandas.Series.apply完成转换:
import json import pandas as pd # 将字符串转为Python列表 df['authorships_parsed'] = df['authorships'].apply(json.loads)
步骤2:将每个作者信息拆分为单独行
使用explode方法把列表里的每个字典拆成单独行,实现"每行对应一位作者"的需求:
# 拆分每个作者为单独行 df_exploded = df.explode('authorships_parsed', ignore_index=True)
步骤3:扁平化嵌套的作者信息
用pd.json_normalize处理拆分后的字典,自动提取author嵌套字段(如id、display_name),同时保留authorships_parsed里的其他字段(如author_position、is_corresponding):
# 扁平化嵌套结构 final_df = pd.json_normalize(df_exploded['authorships_parsed'])
完整代码示例
整合所有步骤的完整代码:
import json import pandas as pd # 假设原始DataFrame名为df # 1. 解析字符串为Python对象 df['authorships_parsed'] = df['authorships'].apply(json.loads) # 2. 拆分每个作者为单行 df_exploded = df.explode('authorships_parsed', ignore_index=True) # 3. 扁平化嵌套数据 final_df = pd.json_normalize(df_exploded['authorships_parsed']) # 查看结果 print(final_df.head())
额外说明
- 最终生成的
final_df中,author.id、author.display_name、author.orcid会作为独立列存在,同时保留author_position、is_corresponding等原有字段。 - 如果原始DataFrame包含文章ID、标题等其他列,可通过以下方式保留并合并:
# 保留原DataFrame的其他列,与扁平化后的作者数据合并 final_df = pd.concat([df_exploded.drop('authorships_parsed', axis=1), final_df], axis=1)
内容的提问来源于stack exchange,提问作者Wildcat PhD
相关产品推荐
相关产品推荐

