使用Python爬取Reddit仅获3-4条结果,求解决方案
问题解决:爬取Reddit r/football更多帖子及绕过限制方法
一、为什么只拿到3-4条帖子?
Reddit是动态加载页面,初始仅会加载顶部少量帖子,剩余内容需要通过滚动页面触发加载。你的代码仅获取了页面初始加载的HTML,自然只能拿到少量数据。
二、获取更多帖子的修改方案
用Selenium模拟页面滚动,触发动态加载,等加载足够多内容后再提取HTML。修改后的代码如下:
import time from selenium import webdriver from selenium.webdriver.chrome.service import Service as ChromeService from webdriver_manager.chrome import ChromeDriverManager import pandas as pd from bs4 import BeautifulSoup # 初始化浏览器 driver = webdriver.Chrome(service=ChromeService(ChromeDriverManager().install())) driver.get("https://www.reddit.com/r/football/") # 模拟滚动加载更多帖子,scroll_count可按需调整 scroll_count = 8 for _ in range(scroll_count): # 滚动到页面底部 driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") # 等待页面加载,网络慢可延长等待时间 time.sleep(3) # 获取完整页面HTML html = driver.page_source soup = BeautifulSoup(html, 'html.parser') # 定位帖子元素(注意:Reddit页面样式可能更新,若失效需重新检查class) post_elements = soup.find_all('shreddit-post', class_='block cursor-pointer relative bg-neutral-background focus-within:bg-neutral-background-hover hover:bg-neutral-background-hover xs:rounded-[16px] p-md my-2xs nd:visible') post_data_list = [] for post_elem in post_elements: post_data = {} # 用get()方法避免属性缺失报错 post_data['post_title'] = post_elem.get('post-title', 'N/A') post_data['permalink'] = post_elem.get('permalink', 'N/A') post_data['author'] = post_elem.get('author', 'N/A') timestamp_elem = post_elem.find('time') post_data['timestamp'] = timestamp_elem['datetime'] if timestamp_elem else 'N/A' post_data['score'] = post_elem.get('score', 'N/A') post_data['domain'] = post_elem.get('domain', 'N/A') post_data_list.append(post_data) # 导出为CSV文件 reddit_df = pd.DataFrame(post_data_list) reddit_df.to_csv('reddit_football_posts.csv', index=False) # 关闭浏览器 driver.quit()
代码关键调整:
- 新增滚动逻辑,通过控制滚动次数加载更多内容
- 添加等待时间,确保新内容完全加载
- 用
get()方法替代直接取属性,避免元素属性缺失导致程序崩溃 - 补充了CSV导出和浏览器关闭的步骤
三、Reddit爬取限制及绕过方法
Reddit有反爬机制,频繁爬取会被限制,以下是可行的应对方案:
模拟真实浏览器行为:给Selenium添加自定义User-Agent,禁用自动化检测标记,避免被识别为爬虫:
from selenium.webdriver.chrome.options import Options chrome_options = Options() chrome_options.add_argument("--disable-blink-features=AutomationControlled") chrome_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36") driver = webdriver.Chrome(service=ChromeService(ChromeDriverManager().install()), options=chrome_options)控制爬取频率:不要快速连续滚动或请求,每次滚动后延长等待时间,避免触发频率限制
使用Reddit官方API:这是最合规的方式,注册Reddit开发者账号获取API密钥,用PRAW库(Python Reddit API Wrapper)获取数据,无反爬风险且数据结构更规整。示例代码(需先安装
praw:pip install praw):import praw import pandas as pd # 替换为你的开发者API信息 reddit = praw.Reddit( client_id="你的Client ID", client_secret="你的Client Secret", user_agent="自定义用户代理名称" ) post_data_list = [] # 获取r/football热门帖子,limit控制数量(最多1000) for submission in reddit.subreddit("football").hot(limit=200): post_data = { "post_title": submission.title, "permalink": submission.permalink, "author": submission.author.name if submission.author else "N/A", "timestamp": submission.created_utc, "score": submission.score, "domain": submission.domain } post_data_list.append(post_data) reddit_df = pd.DataFrame(post_data_list) reddit_df.to_csv('reddit_football_api_posts.csv', index=False)使用稳定代理(可选):若频繁爬取被限制IP,可使用正规代理服务,但需避免频繁更换IP地址
注意事项
- Reddit页面结构(如
shreddit-post的class)可能随版本更新变化,若代码失效,需重新用浏览器开发者工具检查元素属性 - 爬取时遵守Reddit的
robots.txt规则,不要过度爬取,避免账号被封禁
内容的提问来源于stack exchange,提问作者Alex D
相关产品推荐
相关产品推荐

