使用BeautifulSoup爬取Reddit评论时间戳无结果求助
问题描述
研究需求下,尝试爬取某Reddit帖子(约700条评论)的评论时间戳,使用BeautifulSoup实现。编写的Python代码如下:
from bs4 import BeautifulSoup import requests import csv url = "https://www.reddit.com/r/NewTubers/comments/1bfhcwz/feedback_friday_post_your_videos_here_if_you_want/" r = requests.get(url) #print(r.status_code) returned 200 soup = BeautifulSoup(r.content, 'html.parser') #lxml didn't work either #print(soup.title) returned the correct title of the HTML page file = open("scraped_timestamps.csv", "w") writer = csv.writer(file) writer.writerow(["TIMESTAMPS"]) timestamps = soup.findAll('a', class_='_3yx4Dn0W3Yunucf5sVJeFU') for timestamp in timestamps: writer.writerow([timestamp.text]) file.close()
观察页面元素时,评论时间戳均属于<a>标签下的类_3yx4Dn0W3Yunucf5sVJeFU,但通过该类定位后未获取到任何内容,尝试lxml解析器也无效。此外,Reddit鼠标悬停时间戳会显示精确到秒的时间,后续也希望能爬取该数据,目前需先解决基础时间戳爬取问题。
原因分析
Reddit的评论内容是动态加载的:requests.get()只能获取页面的静态HTML骨架,评论及时间戳等内容需要通过JavaScript异步加载,因此静态HTML中不存在你指定的_3yx4Dn0W3Yunucf5sVJeFU类元素,导致BeautifulSoup无法定位到目标内容。
解决方案
方法1:使用Reddit官方API(推荐,合规稳定)
Reddit提供了官方API接口,可直接获取评论的原始时间数据(包括精确到秒的时间戳),无需解析HTML,且不会触发反爬机制。
- 安装
praw(Reddit的Python SDK)及依赖:
pip install praw python-dotenv
- 编写代码:
import praw import csv from datetime import datetime from dotenv import load_dotenv import os # 加载环境变量(建议将API信息存放在.env文件中) load_dotenv() # 初始化Reddit客户端 reddit = praw.Reddit( client_id=os.getenv('REDDIT_CLIENT_ID'), client_secret=os.getenv('REDDIT_CLIENT_SECRET'), user_agent=os.getenv('REDDIT_USER_AGENT') ) # 目标帖子URL post_url = "https://www.reddit.com/r/NewTubers/comments/1bfhcwz/feedback_friday_post_your_videos_here_if_you_want/" submission = reddit.submission(url=post_url) # 加载所有评论(默认只加载部分,需调用replace_more()获取全部) submission.comments.replace_more(limit=None) # 写入CSV with open("scraped_timestamps.csv", "w", newline='', encoding='utf-8') as file: writer = csv.writer(file) writer.writerow(["显示时间", "精确时间(UTC)", "精确时间(本地)"]) for comment in submission.comments.list(): # 页面显示的相对时间(如"3 days ago") display_time = comment.time # 精确到秒的UTC时间戳转换为可读格式 utc_time = datetime.utcfromtimestamp(comment.created_utc).strftime('%Y-%m-%d %H:%M:%S') # 转换为本地时间 local_time = datetime.fromtimestamp(comment.created_utc).strftime('%Y-%m-%d %H:%M:%S') writer.writerow([display_time, utc_time, local_time])
说明:
- 需要先在Reddit开发者平台创建应用,获取
client_id、client_secret并设置user_agent; comment.created_utc就是悬停时显示的精确时间戳,可直接转换为可读格式;comment.time会返回和页面一致的相对时间文本。
方法2:使用Selenium渲染动态页面
如果不想使用API,可通过Selenium模拟浏览器加载页面,等待动态内容渲染完成后再解析:
- 安装依赖:
pip install selenium beautifulsoup4 csv
- 编写代码:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup import csv import time # 初始化Chrome浏览器(需下载对应版本的chromedriver) driver = webdriver.Chrome() driver.get("https://www.reddit.com/r/NewTubers/comments/1bfhcwz/feedback_friday_post_your_videos_here_if_you_want/") # 滚动页面加载所有评论 scroll_pause_time = 2 last_height = driver.execute_script("return document.body.scrollHeight") while True: driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") time.sleep(scroll_pause_time) new_height = driver.execute_script("return document.body.scrollHeight") if new_height == last_height: break last_height = new_height # 等待时间戳元素加载完成 WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, "_3yx4Dn0W3Yunucf5sVJeFU")) ) # 获取页面源码并解析 soup = BeautifulSoup(driver.page_source, 'html.parser') driver.quit() # 写入CSV with open("scraped_timestamps.csv", "w", newline='', encoding='utf-8') as file: writer = csv.writer(file) writer.writerow(["显示时间", "精确时间"]) timestamps = soup.find_all('a', class_='_3yx4Dn0W3Yunucf5sVJeFU') for timestamp in timestamps: # 页面显示的相对时间 display_time = timestamp.text # 悬停显示的精确时间存放在title属性中 exact_time = timestamp.get('title') writer.writerow([display_time, exact_time])
说明:
- 需要下载对应浏览器版本的驱动(如Chrome的chromedriver);
- 滚动页面是为了触发所有评论的加载,避免只爬取初始渲染的部分;
- 悬停的精确时间直接存放在
<a>标签的title属性中,无需额外处理。
内容的提问来源于stack exchange,提问作者KNutellaZ
相关产品推荐
相关产品推荐

