Reddit旧版页面爬取股票数据异常:仅获138条记录求助
爬取Old Reddit r/wallstreetbets数据仅获138条记录的问题排查
我正在进行学校项目,需要获取Reddit平台Wallstreetbets、StockMarket子版块的股票相关数据。因新版Reddit API限制严苛,转而爬取旧版Reddit页面,但即便将num_pages_to_scrape设为5000,也仅得到138条记录。怀疑是next_button功能异常或time.sleep(2)参数需调整,但问题仍未解决,附上我的代码请求协助排查:
import requests from bs4 import BeautifulSoup import pandas as pd import time url = "https://old.reddit.com/r/wallstreetbets" headers = {'User-Agent': 'Mozilla/5.0'} data = [] # List to store post data #Set the desired number of pages num_pages_to_scrape = 5000 for counter in range(1, num_pages_to_scrape + 1): page = requests.get(url, headers=headers) soup = BeautifulSoup(page.text, 'html.parser') posts = soup.find_all('div', class_='thing', attrs={'data-domain': 'self.wallstreetbets'}) for post in posts: title = post.find('a', class_='title').text author = post.find('a', class_='author').text comments = post.find('a', class_='comments').text.split()[0] if comments == "comment": comments = 0 likes = post.find("div", class_="score likes").text if likes == "•": likes = "None" # Extract the date information from the HTML date_element = post.find('time', class_='live-timestamp') date = date_element['datetime'] if date_element else "N/A" formatted_date = pd.to_datetime(date, utc=True).strftime('%Y-%m-%d %H:%M:%S') data.append([formatted_date, title, author, comments, likes]) next_button = soup.find("span", class_="next-button") if next_button: next_page_link = next_button.find("a").attrs['href'] url = next_page_link else: break time.sleep(2) #Create df columns = ['Date', 'Title', 'Author', 'Comments', 'Likes'] df = pd.DataFrame(data, columns=columns) #Print the DataFrame df
可能的问题及修复方案
- Old Reddit分页隐性限制:旧版Reddit不会允许无限制翻取历史数据,通常仅能获取最近几百条帖子,这大概率是你只拿到138条记录的核心原因,无法通过代码完全突破。
- 翻页按钮定位不准确:当前用
span.next-button定位翻页元素的方式不稳定,建议改为直接抓取带有rel="next"属性的链接,兼容性更强:next_link = soup.find('a', rel='next') if next_link: url = next_link['href'] else: break - 请求头过于简单:仅设置
User-Agent易被识别为爬虫,补充完整请求头可降低限流概率:headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept-Language': 'en-US,en;q=0.9', 'Referer': 'https://old.reddit.com/' } - 数据过滤范围过窄:代码仅抓取
data-domain='self.wallstreetbets'的原创帖,若要扩大数据量,可移除该过滤条件,后续再通过标题筛选股票相关内容:posts = soup.find_all('div', class_='thing') - 缺失异常处理机制:请求页面时可能遇到429限流、503服务器错误等问题,添加异常捕获可避免爬取中途中断:
try: page = requests.get(url, headers=headers, timeout=10) page.raise_for_status() # 触发HTTP错误异常 except requests.exceptions.RequestException as e: print(f"第{counter}页请求失败: {e}") time.sleep(5) continue - 请求间隔需优化:固定2秒间隔易被识别为爬虫,建议改为随机间隔(2-5秒):
import random time.sleep(random.uniform(2, 5))
内容的提问来源于stack exchange,提问作者TMO995
相关产品推荐
相关产品推荐

