如何使用PRAW获取Reddit子版块的帖子及对应评论并存储为JSON?
PRAW获取子版块投稿对应评论的实现方法
核心逻辑说明
通过PRAW调取的每一条投稿(submission)对象都自带comments属性,可直接获取该帖子下的评论内容。默认仅返回一级评论,嵌套回复需要调用replace_more方法展开全量层级。
完整实现步骤
1. 全量评论拉取代码实现
import praw import json import time from datetime import datetime # 替换为你自己的Reddit API凭证初始化实例 reddit = praw.Reddit( client_id="你的客户端ID", client_secret="你的客户端密钥", user_agent="你的应用标识" ) all_result = [] # 2015年1月1日的时间戳,用于过滤投稿 start_timestamp = datetime(2015, 1, 1).timestamp() # 可替换为new()/hot()等你需要的分类,limit=None表示拉取所有可访问的投稿 for submission in reddit.subreddit("dogs").top("all", limit=None): # 跳过2015年之前发布的投稿 if submission.created_utc < start_timestamp: continue # 初始化单条投稿的基础信息 post_item = { "post_id": submission.id, "title": submission.title, "author": str(submission.author), "publish_time": datetime.fromtimestamp(submission.created_utc).strftime("%Y-%m-%d %H:%M:%S"), "upvote_num": submission.score, "comment_list": [] } # 展开所有隐藏的嵌套评论,limit=None表示加载全部,可修改数值限制评论加载量减少耗时 submission.comments.replace_more(limit=None) # 遍历所有层级的评论 for comment in submission.comments.list(): comment_item = { "comment_id": comment.id, "comment_author": str(comment.author), "content": comment.body, "comment_publish_time": datetime.fromtimestamp(comment.created_utc).strftime("%Y-%m-%d %H:%M:%S"), "comment_upvote": comment.score, "parent_id": comment.parent_id } post_item["comment_list"].append(comment_item) all_result.append(post_item) # 可选:拉取单条投稿后加延时,避免触发API请求频率限制 time.sleep(0.5)
2. 存储为JSON格式
数据拉取完成后直接用json库写入本地文件即可:
with open("r_dogs_2015_comments.json", "w", encoding="utf-8") as f: json.dump(all_result, f, ensure_ascii=False, indent=4)
3. 注意事项
- 如果不需要全量嵌套评论,可将
replace_more的limit参数修改为你需要的最大加载数量,大幅降低请求耗时 - 如需同时过滤2015年之前发布的评论,可在遍历comment时新增时间戳判断逻辑
- 拉取数据量超过10万条时,建议适当调高延时参数,避免触发API临时封禁规则
内容的提问来源于stack exchange,提问作者ml_learner
相关产品推荐
相关产品推荐

