如何用PRAW批量抓取Reddit上4000条含图帖子及点赞数据?
我正在开展一个数据科学项目,需从Reddit抓取带图片的帖子及其对应点赞数,目标获取3000-4000条样本。但PRAW单次仅支持获取1000条提交内容,我尝试分批抓取:先获取最新1000条,再依次往前抓取更多批次直至满足数量需求,却发现subreddit.new()并不存在预期的before参数。以下是我当前尝试运行的代码:
download_sub = reddit.subreddit(sub_name) MAX_POSTS = 4000 # Maximum number of posts to download post_count = 0 # Counter for number of posts downloaded # get the latest 1000 posts in the subreddit latest_posts = download_sub.new(limit=1000) #latest_posts = list(download_sub.new(limit=1000)) # create an empty dataframe with the specified columns df = pd.DataFrame(columns=['Id','PostedAt', 'Link', 'Username', 'Upvotes']) for post in latest_posts: # get the post creation time in UTC and convert it to a datetime object utc_time = datetime.fromtimestamp(post.created_utc, timezone.utc) created_at = utc_time.strftime('%Y-%m-%d %H:%M:%S') # get the post link, username, image link (if exists), and upvotes link = f"https://www.reddit.com{post.permalink}" username = post.author.name if post.author else '[deleted]' upvotes = post.score # if the post is a gallery, download the gallery, else download the image if hasattr(post, "is_gallery"): df_filename = downloadGallery(post) df = df.append({'Id': post.id, 'PostedAt': created_at, 'Link': link, 'Username': username, 'Filename': df_filename, 'Upvotes': upvotes}, ignore_index=True) elif '.png' in post.url or '.jpg' in post.url or '.jpeg' in post.url: df_filename = downloadImg(post) df = df.append({'Id': post.id, 'PostedAt': created_at, 'Link': link, 'Username': username, 'Filename': df_filename, 'Upvotes': upvotes}, ignore_index=True) os.system('clear') post_count += 1 oldest_post_id = post.id print(post_count, " done so far.") if post_count >= MAX_POSTS: break # get the next 1000 posts before the oldest post in the latest 1000 posts next_posts = download_sub.new(limit=1000, before=oldest_post_id) # loop to get the next 1000 posts in batches until you reach MAX_POSTS while len(next_posts) > 0: for post in next_posts: # get the post creation time in UTC and convert it to a datetime object utc_time = datetime.fromtimestamp(post.created_utc, timezone.utc) created_at = utc_time.strftime('%Y-%m-%d %H:%M:%S') # get the post link, username, image link (if exists), and upvotes link = f"https://www.reddit.com{post.permalink}" username = post.author.name if post.author else '[deleted]' upvotes = post.score # if the post is a gallery, download the gallery, else download the image if hasattr(post, "is_gallery"): df_filename = downloadGallery(post) df = df.append({'Id': post.id, 'PostedAt': created_at, 'Link': link, 'Username': username, 'Filename': df_filename, 'Upvotes': upvotes}, ignore_index=True) elif '.png' in post.url or '.jpg' in post.url or '.jpeg' in post.url: df_filename = downloadImg(post) df = df.append({'Id': post.id, 'PostedAt': created_at, 'Link': link, 'Username': username, 'Filename': df_filename, 'Upvotes': upvotes}, ignore_index=True) os.system('clear') post_count += 1 oldest_post_id = post.id print(post_count, " done so far.") if post_count >= MAX_POSTS: break if post_count >= MAX_POSTS: break # get the next 1000 posts before the oldest post in the next 1000 posts next_posts = download_sub.new(limit=1000, before=oldest_post_id)请问有没有可实现类似分批抓取功能的方法?
解决方案
PRAW的ListingGenerator(即subreddit.new()返回的对象)支持两种方式实现分页抓取,无需手动构造无效参数,以下是具体方法:
方法一:利用ListingGenerator自动分页(推荐)
ListingGenerator本身支持迭代超过1000条数据,只要持续迭代,它会自动向Reddit API请求下一页数据,无需手动处理分页参数。
修改后的完整代码:
import praw import pandas as pd from datetime import datetime, timezone import os # 假设已完成reddit实例初始化 download_sub = reddit.subreddit(sub_name) MAX_POSTS = 4000 post_count = 0 # 修正列名,添加原代码中用到的Filename字段 df = pd.DataFrame(columns=['Id','PostedAt', 'Link', 'Username', 'Filename', 'Upvotes']) # 用limit=None获取所有可迭代的帖子,ListingGenerator自动处理分页 for post in download_sub.new(limit=None): if post_count >= MAX_POSTS: break # 处理帖子时间 utc_time = datetime.fromtimestamp(post.created_utc, timezone.utc) created_at = utc_time.strftime('%Y-%m-%d %H:%M:%S') link = f"https://www.reddit.com{post.permalink}" username = post.author.name if post.author else '[deleted]' upvotes = post.score df_filename = None # 过滤并下载带图片的帖子 if hasattr(post, "is_gallery") and post.is_gallery: df_filename = downloadGallery(post) elif any(ext in post.url for ext in ['.png', '.jpg', '.jpeg']): df_filename = downloadImg(post) # 仅将成功获取图片的帖子加入数据框 if df_filename: # 用loc替代已废弃的append方法,提升效率 df.loc[len(df)] = { 'Id': post.id, 'PostedAt': created_at, 'Link': link, 'Username': username, 'Filename': df_filename, 'Upvotes': upvotes } post_count += 1 os.system('clear') print(f"{post_count} done so far.") # 保存结果 df.to_csv('reddit_image_posts.csv', index=False)
方法二:手动控制分批抓取
如果需要明确控制每批抓取的数量,可以通过params参数传递before参数,注意传入的值必须是帖子的fullname(格式为t3_帖子id),而非单纯的帖子id。
示例代码:
def fetch_batch(subreddit, before_fullname=None): # 构造分页参数 params = {'before': before_fullname} if before_fullname else None return list(subreddit.new(limit=1000, params=params)) post_count = 0 current_before = None df = pd.DataFrame(columns=['Id','PostedAt', 'Link', 'Username', 'Filename', 'Upvotes']) while post_count < MAX_POSTS: batch = fetch_batch(download_sub, current_before) if not batch: break # 无更多帖子可抓取 for post in batch: if post_count >= MAX_POSTS: break # 帖子处理逻辑同方法一,省略重复代码 # ... if df_filename: df.loc[len(df)] = { ... } post_count += 1 print(f"{post_count} done so far.") # 更新下一批的before参数为当前批次最后一个帖子的fullname current_before = batch[-1].fullname
关键注意事项
- 确保使用最新版本的PRAW,避免兼容性问题。
- Reddit API有速率限制,不要过度频繁请求,PRAW默认会处理速率限制,但如果自定义请求逻辑需要注意。
- 原代码中
df.append()已被pandas废弃,建议使用df.loc[len(df)]替代。
内容的提问来源于stack exchange,提问作者erand89
相关产品推荐
相关产品推荐

