如何将DataFrame中含JSON字符串的列拆分至独立列
拆分DataFrame中JSON字符串列生成独立列的实现方法
方法一:使用pd.json_normalize(推荐)
这是处理JSON列最简洁的方式,尤其适合嵌套结构的JSON数据:
import pandas as pd import json # 原始数据 raw_data = [{'id': 1, 'name': 'FRANK', 'attributes': '{"deleted": false, "rejected": true, "handled": true, "order": "37"}'}] raw_df = pd.DataFrame(raw_data) # 将JSON字符串解析为字典格式 raw_df['attributes'] = raw_df['attributes'].apply(json.loads) # 展开JSON列,并与原DataFrame合并 new_df = pd.concat([ raw_df.drop('attributes', axis=1), pd.json_normalize(raw_df['attributes']) ], axis=1) # 将order列转换为数值类型(匹配示例结果) new_df['order'] = pd.to_numeric(new_df['order'])
方法二:使用apply + pd.Series
对于简单的键值对JSON,也可以用这种方式展开:
import pandas as pd import json # 原始数据 raw_data = [{'id': 1, 'name': 'FRANK', 'attributes': '{"deleted": false, "rejected": true, "handled": true, "order": "37"}'}] raw_df = pd.DataFrame(raw_data) # 解析JSON字符串并展开为单独列 attributes_cols = raw_df['attributes'].apply(lambda x: pd.Series(json.loads(x))) # 合并原数据与展开后的列 new_df = pd.concat([raw_df.drop('attributes', axis=1), attributes_cols], axis=1) # 转换order列类型 new_df['order'] = pd.to_numeric(new_df['order'])
执行上述任意一种方法后,new_df的结构就会和需求结果一致:
id name deleted rejected handled order 0 1 FRANK false true true 37
内容的提问来源于stack exchange,提问作者Angie
相关产品推荐
相关产品推荐

