You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Reddit旧版页面爬取股票数据异常:仅获138条记录求助

爬取Old Reddit r/wallstreetbets数据仅获138条记录的问题排查

我正在进行学校项目,需要获取Reddit平台Wallstreetbets、StockMarket子版块的股票相关数据。因新版Reddit API限制严苛,转而爬取旧版Reddit页面,但即便将num_pages_to_scrape设为5000,也仅得到138条记录。怀疑是next_button功能异常或time.sleep(2)参数需调整,但问题仍未解决,附上我的代码请求协助排查:

import requests
from bs4 import BeautifulSoup
import pandas as pd
import time

url = "https://old.reddit.com/r/wallstreetbets"
headers = {'User-Agent': 'Mozilla/5.0'}

data = []  # List to store post data

#Set the desired number of pages
num_pages_to_scrape = 5000

for counter in range(1, num_pages_to_scrape + 1):
    page = requests.get(url, headers=headers)
    soup = BeautifulSoup(page.text, 'html.parser')

    posts = soup.find_all('div', class_='thing', attrs={'data-domain': 'self.wallstreetbets'})

    for post in posts:
        title = post.find('a', class_='title').text
        author = post.find('a', class_='author').text
        comments = post.find('a', class_='comments').text.split()[0]

        if comments == "comment":
            comments = 0

        likes = post.find("div", class_="score likes").text

        if likes == "•":
            likes = "None"
        
        # Extract the date information from the HTML
        date_element = post.find('time', class_='live-timestamp')
        date = date_element['datetime'] if date_element else "N/A"
        formatted_date = pd.to_datetime(date, utc=True).strftime('%Y-%m-%d %H:%M:%S')

        data.append([formatted_date, title, author, comments, likes])

    next_button = soup.find("span", class_="next-button")
    if next_button:
        next_page_link = next_button.find("a").attrs['href']
        url = next_page_link
    else:
        break

    time.sleep(2)

#Create df
columns = ['Date', 'Title', 'Author', 'Comments', 'Likes']
df = pd.DataFrame(data, columns=columns)

#Print the DataFrame
df

可能的问题及修复方案

  • Old Reddit分页隐性限制:旧版Reddit不会允许无限制翻取历史数据,通常仅能获取最近几百条帖子,这大概率是你只拿到138条记录的核心原因,无法通过代码完全突破。
  • 翻页按钮定位不准确:当前用span.next-button定位翻页元素的方式不稳定,建议改为直接抓取带有rel="next"属性的链接,兼容性更强:
    next_link = soup.find('a', rel='next')
    if next_link:
        url = next_link['href']
    else:
        break
    
  • 请求头过于简单:仅设置User-Agent易被识别为爬虫,补充完整请求头可降低限流概率:
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
        'Accept-Language': 'en-US,en;q=0.9',
        'Referer': 'https://old.reddit.com/'
    }
    
  • 数据过滤范围过窄:代码仅抓取data-domain='self.wallstreetbets'的原创帖,若要扩大数据量,可移除该过滤条件,后续再通过标题筛选股票相关内容:
    posts = soup.find_all('div', class_='thing')
    
  • 缺失异常处理机制:请求页面时可能遇到429限流、503服务器错误等问题,添加异常捕获可避免爬取中途中断:
    try:
        page = requests.get(url, headers=headers, timeout=10)
        page.raise_for_status()  # 触发HTTP错误异常
    except requests.exceptions.RequestException as e:
        print(f"第{counter}页请求失败: {e}")
        time.sleep(5)
        continue
    
  • 请求间隔需优化:固定2秒间隔易被识别为爬虫,建议改为随机间隔(2-5秒):
    import random
    time.sleep(random.uniform(2, 5))
    

内容的提问来源于stack exchange,提问作者TMO995

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 16:53:11