You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将双层嵌套Facebook JSON数据转为Pandas DataFrame

解决Facebook嵌套JSON转Pandas DataFrame问题

问题背景

我需要将获取到的Facebook业务数据转换成指定格式的Pandas DataFrame,部分帖子包含评论,部分无评论。期望的DataFrame格式如下:

User ID and Post ID   comment                 Commenter_id
0  user_id_post_id       0 or N/A                  0 or N/A        
1  user_id_post_id1      0 or N/A                  0 or N/A
2  user_id_post_id2      0 or N/A                  0 or N/A
3  user_id_post_id3      Comment_id                user_who_commented_the_id_comment_id
4  user_id_post_id3      Comment_id*               user_who_commented_the_id_comment_id
5  user_id_post_id4      0 or N/A                  0 or N/A

注:Comment_id*表示同一帖子下的另一条评论

现有JSON数据结构示例:

{'data': [{'id': 'user_id_post_id1'},
  {'id': 'user_id_post_id2'},
  {'id': 'user_id_post_id3'},
  {'comments': {'data': [{'created_time': '2022-11-09T00:15:29+0000',
      'message': 'comment_id',
      'id': 'user_who_commented_the_id_comment_id'}]},
   'id': 'user_id_post_id4'},
  {'id': 'user_id_post_id5'}...]}

尝试执行以下代码时触发错误:

df = pd.json_normalize(data=JSON_Name["data"]["comments"])

错误信息:

---------------------------------------------------------------------------
TypeError                                 Traceback (most recent call last)
/tmp/ipykernel_1/FileName.py in <module>
----> 1 df = pd.json_normalize(data=basic_insight["data"]["comments"])

TypeError: list indices must be integers or slices, not str

错误原因

JSON_Name["data"]是列表类型,不能直接通过字符串["comments"]索引列表元素。同时,并非所有帖子都包含comments字段,直接访问会引发KeyError,且原代码未处理无评论的帖子场景。

解决方案

遍历每条帖子数据,针对有无评论的情况分别处理,生成符合要求的字典列表后再转换为DataFrame:

import pandas as pd

# 假设你的JSON数据存储在facebook_data变量中
facebook_data = {'data': [{'id': 'user_id_post_id1'},
  {'id': 'user_id_post_id2'},
  {'id': 'user_id_post_id3'},
  {'comments': {'data': [{'created_time': '2022-11-09T00:15:29+0000',
      'message': 'comment_id',
      'id': 'user_who_commented_the_id_comment_id'},
     {'created_time': '2022-11-10T08:30:00+0000',
      'message': 'comment_id*',
      'id': 'another_commenter_id'}]},
   'id': 'user_id_post_id4'},
  {'id': 'user_id_post_id5'}]}

# 初始化空列表存储处理后的数据
processed_data = []

for post in facebook_data['data']:
    post_id = post['id']
    # 检查当前帖子是否存在有效评论
    if 'comments' in post and 'data' in post['comments'] and len(post['comments']['data']) > 0:
        # 遍历每条评论,生成对应行数据
        for comment in post['comments']['data']:
            processed_data.append({
                'User ID and Post ID': post_id,
                'comment': comment['message'],
                'Commenter_id': comment['id']
            })
    else:
        # 无评论时填充默认值
        processed_data.append({
            'User ID and Post ID': post_id,
            'comment': 'N/A',
            'Commenter_id': 'N/A'
        })

# 转换为目标DataFrame
df = pd.DataFrame(processed_data)
print(df)

执行后输出的DataFrame符合需求:

User ID and Post ID      comment                  Commenter_id
0    user_id_post_id1          N/A                          N/A
1    user_id_post_id2          N/A                          N/A
2    user_id_post_id3          N/A                          N/A
3    user_id_post_id4    comment_id  user_who_commented_the_id_comment_id
4    user_id_post_id4  comment_id*                another_commenter_id
5    user_id_post_id5          N/A                          N/A

内容的提问来源于stack exchange,提问作者Ursa Major

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 02:25:46