如何从字典列表结构的Facebook评论中提取comment_text内容?
提取Facebook评论及回复中的comment_text字段
方法一:爬取阶段直接提取
修改现有爬取代码,在遍历评论时直接提取comment_text,避免存储冗余数据:
import pandas as pd import facebook_scraper post_ids = ['1014199301965488'] options = {"comments": True, "reactors": True, "allow_extra_requests": True, } cookies = "/content/cookies.txt" all_comment_texts = [] for post in facebook_scraper.get_posts(post_urls=post_ids, cookies=cookies, options=options): # 遍历主评论 for comment in post['comments_full']: # 提取主评论文本 all_comment_texts.append(comment['comment_text']) # 检查并遍历回复(facebook_scraper中回复默认存于'comments'键下) if 'comments' in comment and comment['comments']: for reply in comment['comments']: all_comment_texts.append(reply['comment_text']) # 查看提取结果 for text in all_comment_texts: print(text)
方法二:处理已收集的replies列表
如果已经完成爬取并存储了replies列表,直接遍历嵌套结构提取目标字段:
all_comment_texts = [] # 遍历每组主评论+回复的列表 for comment_group in replies: # 遍历组内每个评论/回复字典 for item in comment_group: all_comment_texts.append(item['comment_text']) # 输出所有提取的文本 for text in all_comment_texts: print(text)
补充说明
- 第一种方法更高效,无需存储完整评论字典,仅保留所需文本字段。
- facebook_scraper返回的
comments_full结构中,主评论的回复默认嵌套在comments子键下,需额外遍历该子列表获取回复内容。
内容的提问来源于stack exchange,提问作者PM92
相关产品推荐
相关产品推荐

