You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python爬取Reddit仅获3-4条结果,求解决方案

问题解决:爬取Reddit r/football更多帖子及绕过限制方法

一、为什么只拿到3-4条帖子?

Reddit是动态加载页面,初始仅会加载顶部少量帖子,剩余内容需要通过滚动页面触发加载。你的代码仅获取了页面初始加载的HTML,自然只能拿到少量数据。

二、获取更多帖子的修改方案

用Selenium模拟页面滚动,触发动态加载,等加载足够多内容后再提取HTML。修改后的代码如下:

import time
from selenium import webdriver
from selenium.webdriver.chrome.service import Service as ChromeService
from webdriver_manager.chrome import ChromeDriverManager
import pandas as pd
from bs4 import BeautifulSoup

# 初始化浏览器
driver = webdriver.Chrome(service=ChromeService(ChromeDriverManager().install()))
driver.get("https://www.reddit.com/r/football/")

# 模拟滚动加载更多帖子,scroll_count可按需调整
scroll_count = 8
for _ in range(scroll_count):
    # 滚动到页面底部
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    # 等待页面加载,网络慢可延长等待时间
    time.sleep(3)

# 获取完整页面HTML
html = driver.page_source
soup = BeautifulSoup(html, 'html.parser')

# 定位帖子元素(注意:Reddit页面样式可能更新,若失效需重新检查class)
post_elements = soup.find_all('shreddit-post', class_='block cursor-pointer relative bg-neutral-background focus-within:bg-neutral-background-hover hover:bg-neutral-background-hover xs:rounded-[16px] p-md my-2xs nd:visible')

post_data_list = []

for post_elem in post_elements:
    post_data = {} 
    # 用get()方法避免属性缺失报错
    post_data['post_title'] = post_elem.get('post-title', 'N/A')
    post_data['permalink'] = post_elem.get('permalink', 'N/A')
    post_data['author'] = post_elem.get('author', 'N/A')
    timestamp_elem = post_elem.find('time')
    post_data['timestamp'] = timestamp_elem['datetime'] if timestamp_elem else 'N/A'
    post_data['score'] = post_elem.get('score', 'N/A')
    post_data['domain'] = post_elem.get('domain', 'N/A')
    
    post_data_list.append(post_data)

# 导出为CSV文件
reddit_df = pd.DataFrame(post_data_list)
reddit_df.to_csv('reddit_football_posts.csv', index=False)

# 关闭浏览器
driver.quit()

代码关键调整:

  • 新增滚动逻辑,通过控制滚动次数加载更多内容
  • 添加等待时间,确保新内容完全加载
  • 用get()方法替代直接取属性,避免元素属性缺失导致程序崩溃
  • 补充了CSV导出和浏览器关闭的步骤

三、Reddit爬取限制及绕过方法

Reddit有反爬机制,频繁爬取会被限制,以下是可行的应对方案:

  • 模拟真实浏览器行为:给Selenium添加自定义User-Agent,禁用自动化检测标记,避免被识别为爬虫:

    from selenium.webdriver.chrome.options import Options
    
    chrome_options = Options()
    chrome_options.add_argument("--disable-blink-features=AutomationControlled")
    chrome_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36")
    driver = webdriver.Chrome(service=ChromeService(ChromeDriverManager().install()), options=chrome_options)
    
  • 控制爬取频率:不要快速连续滚动或请求,每次滚动后延长等待时间,避免触发频率限制

  • 使用Reddit官方API:这是最合规的方式,注册Reddit开发者账号获取API密钥,用PRAW库(Python Reddit API Wrapper)获取数据,无反爬风险且数据结构更规整。示例代码(需先安装praw:pip install praw):

    import praw
    import pandas as pd
    
    # 替换为你的开发者API信息
    reddit = praw.Reddit(
        client_id="你的Client ID",
        client_secret="你的Client Secret",
        user_agent="自定义用户代理名称"
    )
    
    post_data_list = []
    # 获取r/football热门帖子,limit控制数量(最多1000)
    for submission in reddit.subreddit("football").hot(limit=200):
        post_data = {
            "post_title": submission.title,
            "permalink": submission.permalink,
            "author": submission.author.name if submission.author else "N/A",
            "timestamp": submission.created_utc,
            "score": submission.score,
            "domain": submission.domain
        }
        post_data_list.append(post_data)
    
    reddit_df = pd.DataFrame(post_data_list)
    reddit_df.to_csv('reddit_football_api_posts.csv', index=False)
    
  • 使用稳定代理(可选):若频繁爬取被限制IP,可使用正规代理服务,但需避免频繁更换IP地址

注意事项

  • Reddit页面结构(如shreddit-post的class)可能随版本更新变化,若代码失效,需重新用浏览器开发者工具检查元素属性
  • 爬取时遵守Reddit的robots.txt规则,不要过度爬取,避免账号被封禁

内容的提问来源于stack exchange,提问作者Alex D

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 10:25:37