如何将双层嵌套Facebook JSON数据转为Pandas DataFrame
解决Facebook嵌套JSON转Pandas DataFrame问题
问题背景
我需要将获取到的Facebook业务数据转换成指定格式的Pandas DataFrame,部分帖子包含评论,部分无评论。期望的DataFrame格式如下:
User ID and Post ID comment Commenter_id 0 user_id_post_id 0 or N/A 0 or N/A 1 user_id_post_id1 0 or N/A 0 or N/A 2 user_id_post_id2 0 or N/A 0 or N/A 3 user_id_post_id3 Comment_id user_who_commented_the_id_comment_id 4 user_id_post_id3 Comment_id* user_who_commented_the_id_comment_id 5 user_id_post_id4 0 or N/A 0 or N/A
注:Comment_id*表示同一帖子下的另一条评论
现有JSON数据结构示例:
{'data': [{'id': 'user_id_post_id1'}, {'id': 'user_id_post_id2'}, {'id': 'user_id_post_id3'}, {'comments': {'data': [{'created_time': '2022-11-09T00:15:29+0000', 'message': 'comment_id', 'id': 'user_who_commented_the_id_comment_id'}]}, 'id': 'user_id_post_id4'}, {'id': 'user_id_post_id5'}...]}
尝试执行以下代码时触发错误:
df = pd.json_normalize(data=JSON_Name["data"]["comments"])
错误信息:
--------------------------------------------------------------------------- TypeError Traceback (most recent call last) /tmp/ipykernel_1/FileName.py in <module> ----> 1 df = pd.json_normalize(data=basic_insight["data"]["comments"]) TypeError: list indices must be integers or slices, not str
错误原因
JSON_Name["data"]是列表类型,不能直接通过字符串["comments"]索引列表元素。同时,并非所有帖子都包含comments字段,直接访问会引发KeyError,且原代码未处理无评论的帖子场景。
解决方案
遍历每条帖子数据,针对有无评论的情况分别处理,生成符合要求的字典列表后再转换为DataFrame:
import pandas as pd # 假设你的JSON数据存储在facebook_data变量中 facebook_data = {'data': [{'id': 'user_id_post_id1'}, {'id': 'user_id_post_id2'}, {'id': 'user_id_post_id3'}, {'comments': {'data': [{'created_time': '2022-11-09T00:15:29+0000', 'message': 'comment_id', 'id': 'user_who_commented_the_id_comment_id'}, {'created_time': '2022-11-10T08:30:00+0000', 'message': 'comment_id*', 'id': 'another_commenter_id'}]}, 'id': 'user_id_post_id4'}, {'id': 'user_id_post_id5'}]} # 初始化空列表存储处理后的数据 processed_data = [] for post in facebook_data['data']: post_id = post['id'] # 检查当前帖子是否存在有效评论 if 'comments' in post and 'data' in post['comments'] and len(post['comments']['data']) > 0: # 遍历每条评论,生成对应行数据 for comment in post['comments']['data']: processed_data.append({ 'User ID and Post ID': post_id, 'comment': comment['message'], 'Commenter_id': comment['id'] }) else: # 无评论时填充默认值 processed_data.append({ 'User ID and Post ID': post_id, 'comment': 'N/A', 'Commenter_id': 'N/A' }) # 转换为目标DataFrame df = pd.DataFrame(processed_data) print(df)
执行后输出的DataFrame符合需求:
User ID and Post ID comment Commenter_id 0 user_id_post_id1 N/A N/A 1 user_id_post_id2 N/A N/A 2 user_id_post_id3 N/A N/A 3 user_id_post_id4 comment_id user_who_commented_the_id_comment_id 4 user_id_post_id4 comment_id* another_commenter_id 5 user_id_post_id5 N/A N/A
内容的提问来源于stack exchange,提问作者Ursa Major
相关产品推荐
相关产品推荐

