Python如何从DataFrame的comments_full列提取comment_id与comment_text值
从pandas DataFrame的comments_full列提取指定评论字段的实现方法
首先明确数据结构:你使用facebook_scraper获取的comments_full列,每个单元格存储的是单条帖子下所有评论组成的字典列表,每个字典自带comment_id、comment_text、comment_time等内置字段,可按以下两种方案提取所需字段:
方法1:展开为单条评论对应一行的结构化表
适合需要单独分析评论、关联对应所属帖子id的场景:
# 先初始化原表(你原有代码漏了这步,直接append会报错) post_df_full = pd.DataFrame() # 原有采集逻辑保持不变,替换已弃用的append方法为concat for post in get_posts('page_name', cookies='C:/Users/stefa/OneDrive/Bureau/cookies.txt', extra_info=True, pages=2, options={"comments": True}): post_entry = post fb_post_df = pd.DataFrame.from_dict(post_entry, orient='index') fb_post_df = fb_post_df.transpose() post_df_full = pd.concat([post_df_full, fb_post_df], ignore_index=True) print(post['post_id']+' get') # 提取评论字段逻辑 # 1. 将评论列表展开,每条评论单独占一行 exploded_post_df = post_df_full[["post_id", "comments_full"]].explode("comments_full", ignore_index=True) # 2. 解析评论字典,提取指定字段 comment_result = pd.json_normalize(exploded_post_df["comments_full"])[["comment_id", "comment_text"]] # 3. 可选:关联对应所属帖子的post_id comment_result = pd.concat([exploded_post_df[["post_id"]], comment_result], axis=1) # 查看结果 comment_result.head()
方法2:保留原表结构,新增列存储评论字段列表
适合需要保留每帖一行的结构,把同帖的所有评论id、文本存在列表里的场景:
# 提取comment_id列表 post_df_full["comment_id_list"] = post_df_full["comments_full"].apply( lambda x: [c["comment_id"] for c in x] if x else [] ) # 提取comment_text列表 post_df_full["comment_text_list"] = post_df_full["comments_full"].apply( lambda x: [c["comment_text"] for c in x] if x else [] ) # 查看结果 post_df_full[["post_id", "comment_id_list", "comment_text_list"]].head()
内容的提问来源于stack exchange,提问作者mohammed chaaraoui
相关产品推荐
相关产品推荐

