You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python:Reddit PRAW拉取大线程评论过慢,API直连结果不全

构建Reddit评论情感机器人:评论爬取的完整性与效率问题

我要构建一个基于Reddit评论的情感机器人,使用Python的PRAW库时发现,获取线程全部评论存在难度:PRAW的代码逻辑简洁,但拉取大线程评论时速度极慢。

于是我参考相关方案实现了直接调用Reddit JSON API的爬取方法,但返回的评论数量和PRAW方法不一致,希望找到缺失评论的原因,实现100%的爬取完整性,同时保持高效的爬取速度。

PRAW方法

import praw

reddit = praw.Reddit(client_id="<MYKEY>",
                     client_secret="<MY_SECRET_KEY>",
                     user_agent="USERAGENT",
                     check_for_async=False)

url = "https://www.reddit.com/r/CryptoCurrency/comments/11rfcjy/daily_general_discussion_march_15_2023_gmt0/"
submission = reddit.submission(url=url)
submission.comments.replace_more(limit=None)
comments = submission.comments.list()
print("top level comments:", submission.comments.__len__())
print("total comments:", len(comments))

JSON API方法

import requests
import time
import numpy as np
import random

# 获取token的方法参考相关文档
TOKEN = get_reddit_token()

headers = {"user-agent": "Mozilla/5.0"}
url = "https://www.reddit.com/r/CryptoCurrency/comments/11rfcjy/daily_general_discussion_march_15_2023_gmt0/.json"
req = requests.get(url, headers=headers)
res = req.json()
body = res[1]["data"]["children"]
thread_id = body[-1]["data"]["parent_id"]

all_comments = {c["data"]["id"]: c for c in body if c["kind"] == "t1"}
comment_ids = [c["data"]["id"] for c in body if c["kind"] == "t1"]
comment_ids += body[-1]["data"]["children"]

def get_more_comments(more_ids, get_replies=True):

    more_children = ",".join(more_ids)
    more_url = f"https://oauth.reddit.com/api/morechildren/.json?api_type=json&link_id={thread_id}&children={more_children}&sort=top"
    
    headers = {'user-agent': 'USERAGENT', 'Authorization': f"bearer {TOKEN}"}
    res = requests.get(more_url, headers=headers) 
    comments = res.json()
    comments = comments["json"]["data"]["things"]
    t1_comments = [c for c in comments if c["kind"] == "t1"]
    more_comments = [c for c in comments if c["kind"] == "more"]
    print("new comments", len(t1_comments))
    print("more comments", len(more_comments))
    
    for comment in t1_comments:
        comment_id = comment["data"]["id"]
        all_comments[comment_id] = comment
    
    if get_replies:
        more_comments = [c["data"]["children"] for c in more_comments]
    else:
        more_comments = [c["data"]["children"] for c in more_comments if c["data"]["parent_id"] == thread_id]
    
    # 扁平化列表
    more_comments = [c for c_list in more_comments for c in c_list]
    
    return more_comments
    
for i in range(100):
    print(i)
    existing_comments = list(all_comments.keys())
    eligible_comments = np.isin(comment_ids, existing_comments)
    eligible_comments = np.array(comment_ids)[~eligible_comments].tolist()
    more_ids = eligible_comments[:100]
    more_comments = get_more_comments(more_ids)
    comment_ids += more_comments
    comment_ids = list(set(comment_ids))
    random.shuffle(comment_ids)
    
    time.sleep(1)

print("top level comments:", len([k for k,v in all_comments.items() if v["data"]["depth"] == 0]))
print("total comments:", len(all_comments.keys()))

方法说明与测试结果

JSON方法的核心思路是:获取线程初始JSON响应(URL末尾添加.json),捕获所有kind=more项的评论ID,循环拉取更多评论。测试中设置了固定迭代次数,一段时间后无法获取新评论,但整体速度远快于PRAW。

  • PRAW方法:耗时超30分钟,返回1259条顶级评论、4686条总评论
  • JSON API方法:返回1259条顶级评论(与PRAW一致),但总评论仅4511条,少于PRAW

注:最初JSON方法返回的评论数量差距更大,在请求URL中添加sort=top(sort=new同样有效)后结果已接近PRAW,但仍存在缺失,希望找到原因以实现完整爬取。

内容的提问来源于stack exchange,提问作者SuperCodeBrah

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 12:05:06