You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取Reddit评论时间戳无结果求助

问题描述

研究需求下,尝试爬取某Reddit帖子(约700条评论)的评论时间戳,使用BeautifulSoup实现。编写的Python代码如下:

from bs4 import BeautifulSoup
import requests
import csv

url = "https://www.reddit.com/r/NewTubers/comments/1bfhcwz/feedback_friday_post_your_videos_here_if_you_want/"
r = requests.get(url)
#print(r.status_code) returned 200

soup = BeautifulSoup(r.content, 'html.parser') #lxml didn't work either
#print(soup.title) returned the correct title of the HTML page
file = open("scraped_timestamps.csv", "w")
writer = csv.writer(file)
writer.writerow(["TIMESTAMPS"])

timestamps = soup.findAll('a', class_='_3yx4Dn0W3Yunucf5sVJeFU')
for timestamp in timestamps:
    writer.writerow([timestamp.text])
file.close()

观察页面元素时,评论时间戳均属于<a>标签下的类_3yx4Dn0W3Yunucf5sVJeFU,但通过该类定位后未获取到任何内容,尝试lxml解析器也无效。此外,Reddit鼠标悬停时间戳会显示精确到秒的时间,后续也希望能爬取该数据,目前需先解决基础时间戳爬取问题。

原因分析

Reddit的评论内容是动态加载的:requests.get()只能获取页面的静态HTML骨架,评论及时间戳等内容需要通过JavaScript异步加载,因此静态HTML中不存在你指定的_3yx4Dn0W3Yunucf5sVJeFU类元素,导致BeautifulSoup无法定位到目标内容。

解决方案

方法1:使用Reddit官方API(推荐,合规稳定)

Reddit提供了官方API接口,可直接获取评论的原始时间数据(包括精确到秒的时间戳),无需解析HTML,且不会触发反爬机制。

  • 安装praw(Reddit的Python SDK)及依赖:
pip install praw python-dotenv
  • 编写代码:
import praw
import csv
from datetime import datetime
from dotenv import load_dotenv
import os

# 加载环境变量(建议将API信息存放在.env文件中)
load_dotenv()

# 初始化Reddit客户端
reddit = praw.Reddit(
    client_id=os.getenv('REDDIT_CLIENT_ID'),
    client_secret=os.getenv('REDDIT_CLIENT_SECRET'),
    user_agent=os.getenv('REDDIT_USER_AGENT')
)

# 目标帖子URL
post_url = "https://www.reddit.com/r/NewTubers/comments/1bfhcwz/feedback_friday_post_your_videos_here_if_you_want/"
submission = reddit.submission(url=post_url)

# 加载所有评论(默认只加载部分,需调用replace_more()获取全部)
submission.comments.replace_more(limit=None)

# 写入CSV
with open("scraped_timestamps.csv", "w", newline='', encoding='utf-8') as file:
    writer = csv.writer(file)
    writer.writerow(["显示时间", "精确时间(UTC)", "精确时间(本地)"])
    
    for comment in submission.comments.list():
        # 页面显示的相对时间(如"3 days ago")
        display_time = comment.time
        # 精确到秒的UTC时间戳转换为可读格式
        utc_time = datetime.utcfromtimestamp(comment.created_utc).strftime('%Y-%m-%d %H:%M:%S')
        # 转换为本地时间
        local_time = datetime.fromtimestamp(comment.created_utc).strftime('%Y-%m-%d %H:%M:%S')
        writer.writerow([display_time, utc_time, local_time])

说明:

  • 需要先在Reddit开发者平台创建应用,获取client_id、client_secret并设置user_agent;
  • comment.created_utc就是悬停时显示的精确时间戳,可直接转换为可读格式;
  • comment.time会返回和页面一致的相对时间文本。

方法2:使用Selenium渲染动态页面

如果不想使用API,可通过Selenium模拟浏览器加载页面,等待动态内容渲染完成后再解析:

  • 安装依赖:
pip install selenium beautifulsoup4 csv
  • 编写代码:
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup
import csv
import time

# 初始化Chrome浏览器(需下载对应版本的chromedriver)
driver = webdriver.Chrome()
driver.get("https://www.reddit.com/r/NewTubers/comments/1bfhcwz/feedback_friday_post_your_videos_here_if_you_want/")

# 滚动页面加载所有评论
scroll_pause_time = 2
last_height = driver.execute_script("return document.body.scrollHeight")

while True:
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    time.sleep(scroll_pause_time)
    new_height = driver.execute_script("return document.body.scrollHeight")
    if new_height == last_height:
        break
    last_height = new_height

# 等待时间戳元素加载完成
WebDriverWait(driver, 10).until(
    EC.presence_of_element_located((By.CLASS_NAME, "_3yx4Dn0W3Yunucf5sVJeFU"))
)

# 获取页面源码并解析
soup = BeautifulSoup(driver.page_source, 'html.parser')
driver.quit()

# 写入CSV
with open("scraped_timestamps.csv", "w", newline='', encoding='utf-8') as file:
    writer = csv.writer(file)
    writer.writerow(["显示时间", "精确时间"])
    
    timestamps = soup.find_all('a', class_='_3yx4Dn0W3Yunucf5sVJeFU')
    for timestamp in timestamps:
        # 页面显示的相对时间
        display_time = timestamp.text
        # 悬停显示的精确时间存放在title属性中
        exact_time = timestamp.get('title')
        writer.writerow([display_time, exact_time])

说明:

  • 需要下载对应浏览器版本的驱动(如Chrome的chromedriver);
  • 滚动页面是为了触发所有评论的加载,避免只爬取初始渲染的部分;
  • 悬停的精确时间直接存放在<a>标签的title属性中,无需额外处理。

内容的提问来源于stack exchange,提问作者KNutellaZ

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 20:59:51