You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过Reddit JSON API获取评论回复(不使用PRAW)

解决方法

核心逻辑

Reddit评论是嵌套树状结构,且存在more类型节点隐藏未加载的评论。要完整获取所有层级,需递归遍历评论树+处理more节点补全隐藏评论,无需依赖PRAW。

具体实现步骤

  1. 配置基础请求参数
    必须设置合法的User-Agent(比如my-comment-fetcher/1.0 by your_username),否则会被API拦截。

  2. 递归遍历评论树函数
    写一个递归函数,处理单条评论及其所有层级回复,同时处理more节点:

    import requests
    import pandas as pd
    
    def parse_comment_tree(comment_node, all_comments):
        # 提取当前评论的核心字段(根据需求调整)
        if comment_node['kind'] == 't1':
            comment_data = {
                'comment_id': comment_node['data']['id'],
                'parent_id': comment_node['data']['parent_id'],
                'author': comment_node['data']['author'],
                'body': comment_node['data']['body'],
                'score': comment_node['data']['score']
            }
            all_comments.append(comment_data)
    
            # 处理当前评论的直接回复
            if comment_node['data']['replies']:
                replies = comment_node['data']['replies']['data']['children']
                for reply in replies:
                    if reply['kind'] == 't1':
                        parse_comment_tree(reply, all_comments)
                    elif reply['kind'] == 'more':
                        # 处理more节点,获取隐藏的评论
                        more_ids = ','.join(reply['data']['children'])
                        more_url = f"https://www.reddit.com/api/morechildren.json?children={more_ids}&api_type=json"
                        headers = {'User-Agent': 'my-comment-fetcher/1.0 by your_username'}
                        more_response = requests.get(more_url, headers=headers)
                        more_data = more_response.json()
                        # 遍历返回的more评论,继续递归处理
                        for more_comment in more_data['json']['data']['things']:
                            parse_comment_tree(more_comment, all_comments)
    
    # 示例:获取单条热门帖子的所有评论
    def fetch_post_all_comments(post_id):
        headers = {'User-Agent': 'my-comment-fetcher/1.0 by your_username'}
        # 初始请求帖子评论,depth=∞确保层级不被截断
        post_url = f"https://www.reddit.com/r/wallstreetbets/comments/{post_id}.json?depth=∞"
        response = requests.get(post_url, headers=headers)
        post_data = response.json()
        
        all_comments = []
        # 第一个元素是帖子信息,第二个是评论列表
        comment_list = post_data[1]['data']['children']
        for comment in comment_list:
            parse_comment_tree(comment, all_comments)
        
        # 转为DataFrame
        return pd.DataFrame(all_comments)
    
  3. 与现有流程整合
    从你已获取的热门帖子DataFrame中提取id字段,循环调用fetch_post_all_comments函数,即可获取每个帖子的完整评论树。

注意事项

  • 速率限制:Reddit API默认限制每分钟60次请求,避免短时间内大量调用导致被封禁。
  • 字段调整:根据实际需求修改comment_data中的字段,比如添加创建时间created_utc等。
  • 异常处理:建议在请求中添加try-except块,处理网络错误、API返回异常等情况。

内容的提问来源于stack exchange,提问作者user19890024

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 03:50:29