You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用PRAW批量抓取Reddit上4000条含图帖子及点赞数据?

我正在开展一个数据科学项目,需从Reddit抓取带图片的帖子及其对应点赞数,目标获取3000-4000条样本。但PRAW单次仅支持获取1000条提交内容,我尝试分批抓取:先获取最新1000条,再依次往前抓取更多批次直至满足数量需求,却发现subreddit.new()并不存在预期的before参数。以下是我当前尝试运行的代码:

download_sub = reddit.subreddit(sub_name)

MAX_POSTS = 4000  # Maximum number of posts to download
post_count = 0    # Counter for number of posts downloaded

# get the latest 1000 posts in the subreddit
latest_posts = download_sub.new(limit=1000)
#latest_posts = list(download_sub.new(limit=1000))


# create an empty dataframe with the specified columns
df = pd.DataFrame(columns=['Id','PostedAt', 'Link', 'Username', 'Upvotes'])

for post in latest_posts:
   # get the post creation time in UTC and convert it to a datetime object
    utc_time = datetime.fromtimestamp(post.created_utc, timezone.utc)
    created_at = utc_time.strftime('%Y-%m-%d %H:%M:%S')

    # get the post link, username, image link (if exists), and upvotes
    link = f"https://www.reddit.com{post.permalink}"
    username = post.author.name if post.author else '[deleted]'
    upvotes = post.score

    # if the post is a gallery, download the gallery, else download the image    

    if hasattr(post, "is_gallery"):
        df_filename = downloadGallery(post)
        df = df.append({'Id': post.id,
                    'PostedAt': created_at,
                    'Link': link,
                    'Username': username,
                    'Filename': df_filename,
                    'Upvotes': upvotes}, ignore_index=True)
    
    elif '.png' in post.url or '.jpg' in post.url or '.jpeg' in post.url:  
        df_filename = downloadImg(post)
        df = df.append({'Id': post.id,
                'PostedAt': created_at,
                'Link': link,
                'Username': username,
                'Filename': df_filename,
                'Upvotes': upvotes}, ignore_index=True)
        
    os.system('clear')
    post_count += 1
    oldest_post_id = post.id
    print(post_count, " done so far.")
    if post_count >= MAX_POSTS:
        break


# get the next 1000 posts before the oldest post in the latest 1000 posts
next_posts = download_sub.new(limit=1000, before=oldest_post_id)

# loop to get the next 1000 posts in batches until you reach MAX_POSTS

while len(next_posts) > 0:
    for post in next_posts:
         # get the post creation time in UTC and convert it to a datetime object
        utc_time = datetime.fromtimestamp(post.created_utc, timezone.utc)
        created_at = utc_time.strftime('%Y-%m-%d %H:%M:%S')

        # get the post link, username, image link (if exists), and upvotes
        link = f"https://www.reddit.com{post.permalink}"
        username = post.author.name if post.author else '[deleted]'
        upvotes = post.score
        
        # if the post is a gallery, download the gallery, else download the image
        
        if hasattr(post, "is_gallery"):
             df_filename = downloadGallery(post)
             df = df.append({'Id': post.id,
                    'PostedAt': created_at,
                    'Link': link,
                    'Username': username,
                    'Filename': df_filename,
                    'Upvotes': upvotes}, ignore_index=True)

        elif '.png' in post.url or '.jpg' in post.url or '.jpeg' in post.url:  
            df_filename = downloadImg(post)
            df = df.append({'Id': post.id,
                'PostedAt': created_at,
                'Link': link,
                'Username': username,
                'Filename': df_filename,
                'Upvotes': upvotes}, ignore_index=True)
            
        os.system('clear')
        post_count += 1
        oldest_post_id = post.id
        print(post_count, " done so far.")
        if post_count >= MAX_POSTS:
            break
        
    if post_count >= MAX_POSTS:
        break
        
    
    # get the next 1000 posts before the oldest post in the next 1000 posts
    next_posts = download_sub.new(limit=1000, before=oldest_post_id)

请问有没有可实现类似分批抓取功能的方法?

解决方案

PRAW的ListingGenerator(即subreddit.new()返回的对象)支持两种方式实现分页抓取,无需手动构造无效参数,以下是具体方法:

方法一:利用ListingGenerator自动分页(推荐)

ListingGenerator本身支持迭代超过1000条数据,只要持续迭代,它会自动向Reddit API请求下一页数据,无需手动处理分页参数。

修改后的完整代码:

import praw
import pandas as pd
from datetime import datetime, timezone
import os

# 假设已完成reddit实例初始化
download_sub = reddit.subreddit(sub_name)

MAX_POSTS = 4000
post_count = 0
# 修正列名,添加原代码中用到的Filename字段
df = pd.DataFrame(columns=['Id','PostedAt', 'Link', 'Username', 'Filename', 'Upvotes'])

# 用limit=None获取所有可迭代的帖子,ListingGenerator自动处理分页
for post in download_sub.new(limit=None):
    if post_count >= MAX_POSTS:
        break
    
    # 处理帖子时间
    utc_time = datetime.fromtimestamp(post.created_utc, timezone.utc)
    created_at = utc_time.strftime('%Y-%m-%d %H:%M:%S')
    
    link = f"https://www.reddit.com{post.permalink}"
    username = post.author.name if post.author else '[deleted]'
    upvotes = post.score
    df_filename = None
    
    # 过滤并下载带图片的帖子
    if hasattr(post, "is_gallery") and post.is_gallery:
        df_filename = downloadGallery(post)
    elif any(ext in post.url for ext in ['.png', '.jpg', '.jpeg']):
        df_filename = downloadImg(post)
    
    # 仅将成功获取图片的帖子加入数据框
    if df_filename:
        # 用loc替代已废弃的append方法,提升效率
        df.loc[len(df)] = {
            'Id': post.id,
            'PostedAt': created_at,
            'Link': link,
            'Username': username,
            'Filename': df_filename,
            'Upvotes': upvotes
        }
        post_count += 1
        os.system('clear')
        print(f"{post_count} done so far.")

# 保存结果
df.to_csv('reddit_image_posts.csv', index=False)

方法二:手动控制分批抓取

如果需要明确控制每批抓取的数量,可以通过params参数传递before参数,注意传入的值必须是帖子的fullname(格式为t3_帖子id),而非单纯的帖子id。

示例代码:

def fetch_batch(subreddit, before_fullname=None):
    # 构造分页参数
    params = {'before': before_fullname} if before_fullname else None
    return list(subreddit.new(limit=1000, params=params))

post_count = 0
current_before = None
df = pd.DataFrame(columns=['Id','PostedAt', 'Link', 'Username', 'Filename', 'Upvotes'])

while post_count < MAX_POSTS:
    batch = fetch_batch(download_sub, current_before)
    if not batch:
        break  # 无更多帖子可抓取
    
    for post in batch:
        if post_count >= MAX_POSTS:
            break
        
        # 帖子处理逻辑同方法一,省略重复代码
        # ...
        
        if df_filename:
            df.loc[len(df)] = { ... }
            post_count += 1
            print(f"{post_count} done so far.")
    
    # 更新下一批的before参数为当前批次最后一个帖子的fullname
    current_before = batch[-1].fullname

关键注意事项

  • 确保使用最新版本的PRAW,避免兼容性问题。
  • Reddit API有速率限制,不要过度频繁请求,PRAW默认会处理速率限制,但如果自定义请求逻辑需要注意。
  • 原代码中df.append()已被pandas废弃,建议使用df.loc[len(df)]替代。

内容的提问来源于stack exchange,提问作者erand89

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 03:25:10