You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python如何从DataFrame的comments_full列提取comment_id与comment_text值

从pandas DataFrame的comments_full列提取指定评论字段的实现方法

首先明确数据结构:你使用facebook_scraper获取的comments_full列,每个单元格存储的是单条帖子下所有评论组成的字典列表,每个字典自带comment_id、comment_text、comment_time等内置字段,可按以下两种方案提取所需字段:

方法1:展开为单条评论对应一行的结构化表

适合需要单独分析评论、关联对应所属帖子id的场景:

# 先初始化原表(你原有代码漏了这步,直接append会报错)
post_df_full = pd.DataFrame()

# 原有采集逻辑保持不变,替换已弃用的append方法为concat
for post in get_posts('page_name', cookies='C:/Users/stefa/OneDrive/Bureau/cookies.txt', extra_info=True, pages=2, options={"comments": True}):
    post_entry = post
    fb_post_df = pd.DataFrame.from_dict(post_entry, orient='index')
    fb_post_df = fb_post_df.transpose()
    post_df_full = pd.concat([post_df_full, fb_post_df], ignore_index=True)
    print(post['post_id']+' get')

# 提取评论字段逻辑
# 1. 将评论列表展开,每条评论单独占一行
exploded_post_df = post_df_full[["post_id", "comments_full"]].explode("comments_full", ignore_index=True)
# 2. 解析评论字典,提取指定字段
comment_result = pd.json_normalize(exploded_post_df["comments_full"])[["comment_id", "comment_text"]]
# 3. 可选:关联对应所属帖子的post_id
comment_result = pd.concat([exploded_post_df[["post_id"]], comment_result], axis=1)

# 查看结果
comment_result.head()

方法2:保留原表结构,新增列存储评论字段列表

适合需要保留每帖一行的结构,把同帖的所有评论id、文本存在列表里的场景:

# 提取comment_id列表
post_df_full["comment_id_list"] = post_df_full["comments_full"].apply(
    lambda x: [c["comment_id"] for c in x] if x else []
)
# 提取comment_text列表
post_df_full["comment_text_list"] = post_df_full["comments_full"].apply(
    lambda x: [c["comment_text"] for c in x] if x else []
)

# 查看结果
post_df_full[["post_id", "comment_id_list", "comment_text_list"]].head()

内容的提问来源于stack exchange,提问作者mohammed chaaraoui

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 23:15:07