You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

新手求助:如何用BeautifulSoup爬取YouTube评论?附代码问题

解决YouTube评论爬取的问题

嘿,我来帮你搞定这个问题!先看看你代码里的几个小错误,再说说为什么拿不到评论的核心原因。

先修正代码里的语法问题

你的代码有几处明显的拼写/语法错误,先改过来:

import requests
from bs4 import BeautifulSoup

# 补上缺失的闭合括号
r = requests.get("https://www.youtube.com/watch?v=kffacxfA7G4")
# 拼写错误:conten → content
req = r.content
soup = BeautifulSoup(req,'html.parser')
print(soup.prettify())
# all是Python内置关键字,换个变量名避免冲突
comment_containers = soup.find_all('div',{'id' : 'contents'})

不过就算修正了这些,你还是拿不到评论——这是因为YouTube的评论是通过JavaScript动态加载的:requests只能获取页面的静态HTML源码,而评论内容是页面加载完成后,通过AJAX请求异步拉取渲染的,所以BeautifulSoup根本看不到这些动态生成的元素。

两种可行的解决方案

方案1:用Selenium模拟浏览器加载(适合快速上手)

Selenium可以模拟真实浏览器的行为,等页面完全加载(包括动态内容)后再抓取源码,这样就能拿到评论了。

首先安装依赖:

pip install selenium

(还要下载对应浏览器的驱动,比如ChromeDriver,记得把路径配置正确)

示例代码:

from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from bs4 import BeautifulSoup
import time

# 初始化Chrome浏览器,替换成你的ChromeDriver路径
service = Service('path/to/chromedriver')
driver = webdriver.Chrome(service=service)

# 打开目标视频页面
driver.get("https://www.youtube.com/watch?v=kffacxfA7G4")

# 等待页面加载,模拟滚动加载更多评论(YouTube需要滚动才会加载更多)
time.sleep(3)  # 先等3秒让基础内容加载
# 滚动3次,每次加载后等2秒
for _ in range(3):
    driver.execute_script("window.scrollTo(0, document.documentElement.scrollHeight);")
    time.sleep(2)

# 获取完全加载后的页面源码
page_source = driver.page_source
soup = BeautifulSoup(page_source, 'html.parser')

# 定位评论容器(YouTube的class可能会更新,记得检查页面结构)
comments = soup.find_all('ytd-comment-thread-renderer', class_='style-scope ytd-comment-section-renderer')

# 提取评论内容
for comment in comments:
    author = comment.find('span', class_='style-scope ytd-comment-renderer').text
    content = comment.find('yt-formatted-string', class_='style-scope ytd-comment-renderer').text
    print(f"作者:{author}\n评论:{content}\n---")

# 关闭浏览器
driver.quit()

方案2:使用YouTube Data API(更稳定可靠)

如果长期爬取评论,官方API是更好的选择——不会因为页面结构变化而失效,还能直接拿到结构化的数据。

步骤大概是:

  1. 去Google Cloud控制台创建项目,启用YouTube Data API v3
  2. 获取你的API密钥
  3. 用requests调用API获取评论

示例代码:

import requests

API_KEY = "你的API密钥"
VIDEO_ID = "kffacxfA7G4"
# 调用API的接口,maxResults可以调整单次获取的评论数量
url = f"https://www.googleapis.com/youtube/v3/commentThreads?part=snippet&videoId={VIDEO_ID}&key={API_KEY}&maxResults=100"

response = requests.get(url)
data = response.json()

# 提取评论内容
for item in data['items']:
    comment_info = item['snippet']['topLevelComment']['snippet']
    author = comment_info['authorDisplayName']
    content = comment_info['textDisplay']
    print(f"作者:{author}\n评论:{content}\n---")

注意:API有免费调用额度,足够个人使用,若需要更多请求可以升级付费套餐。

内容的提问来源于stack exchange,提问作者Abhilash

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 06:47:59