Python:Reddit PRAW拉取大线程评论过慢,API直连结果不全
构建Reddit评论情感机器人:评论爬取的完整性与效率问题
我要构建一个基于Reddit评论的情感机器人,使用Python的PRAW库时发现,获取线程全部评论存在难度:PRAW的代码逻辑简洁,但拉取大线程评论时速度极慢。
于是我参考相关方案实现了直接调用Reddit JSON API的爬取方法,但返回的评论数量和PRAW方法不一致,希望找到缺失评论的原因,实现100%的爬取完整性,同时保持高效的爬取速度。
PRAW方法
import praw reddit = praw.Reddit(client_id="<MYKEY>", client_secret="<MY_SECRET_KEY>", user_agent="USERAGENT", check_for_async=False) url = "https://www.reddit.com/r/CryptoCurrency/comments/11rfcjy/daily_general_discussion_march_15_2023_gmt0/" submission = reddit.submission(url=url) submission.comments.replace_more(limit=None) comments = submission.comments.list() print("top level comments:", submission.comments.__len__()) print("total comments:", len(comments))
JSON API方法
import requests import time import numpy as np import random # 获取token的方法参考相关文档 TOKEN = get_reddit_token() headers = {"user-agent": "Mozilla/5.0"} url = "https://www.reddit.com/r/CryptoCurrency/comments/11rfcjy/daily_general_discussion_march_15_2023_gmt0/.json" req = requests.get(url, headers=headers) res = req.json() body = res[1]["data"]["children"] thread_id = body[-1]["data"]["parent_id"] all_comments = {c["data"]["id"]: c for c in body if c["kind"] == "t1"} comment_ids = [c["data"]["id"] for c in body if c["kind"] == "t1"] comment_ids += body[-1]["data"]["children"] def get_more_comments(more_ids, get_replies=True): more_children = ",".join(more_ids) more_url = f"https://oauth.reddit.com/api/morechildren/.json?api_type=json&link_id={thread_id}&children={more_children}&sort=top" headers = {'user-agent': 'USERAGENT', 'Authorization': f"bearer {TOKEN}"} res = requests.get(more_url, headers=headers) comments = res.json() comments = comments["json"]["data"]["things"] t1_comments = [c for c in comments if c["kind"] == "t1"] more_comments = [c for c in comments if c["kind"] == "more"] print("new comments", len(t1_comments)) print("more comments", len(more_comments)) for comment in t1_comments: comment_id = comment["data"]["id"] all_comments[comment_id] = comment if get_replies: more_comments = [c["data"]["children"] for c in more_comments] else: more_comments = [c["data"]["children"] for c in more_comments if c["data"]["parent_id"] == thread_id] # 扁平化列表 more_comments = [c for c_list in more_comments for c in c_list] return more_comments for i in range(100): print(i) existing_comments = list(all_comments.keys()) eligible_comments = np.isin(comment_ids, existing_comments) eligible_comments = np.array(comment_ids)[~eligible_comments].tolist() more_ids = eligible_comments[:100] more_comments = get_more_comments(more_ids) comment_ids += more_comments comment_ids = list(set(comment_ids)) random.shuffle(comment_ids) time.sleep(1) print("top level comments:", len([k for k,v in all_comments.items() if v["data"]["depth"] == 0])) print("total comments:", len(all_comments.keys()))
方法说明与测试结果
JSON方法的核心思路是:获取线程初始JSON响应(URL末尾添加.json),捕获所有kind=more项的评论ID,循环拉取更多评论。测试中设置了固定迭代次数,一段时间后无法获取新评论,但整体速度远快于PRAW。
- PRAW方法:耗时超30分钟,返回1259条顶级评论、4686条总评论
- JSON API方法:返回1259条顶级评论(与PRAW一致),但总评论仅4511条,少于PRAW
注:最初JSON方法返回的评论数量差距更大,在请求URL中添加sort=top(sort=new同样有效)后结果已接近PRAW,但仍存在缺失,希望找到原因以实现完整爬取。
内容的提问来源于stack exchange,提问作者SuperCodeBrah
相关产品推荐
相关产品推荐

