新手求助:如何用BeautifulSoup爬取YouTube评论?附代码问题
解决YouTube评论爬取的问题
嘿,我来帮你搞定这个问题!先看看你代码里的几个小错误,再说说为什么拿不到评论的核心原因。
先修正代码里的语法问题
你的代码有几处明显的拼写/语法错误,先改过来:
import requests from bs4 import BeautifulSoup # 补上缺失的闭合括号 r = requests.get("https://www.youtube.com/watch?v=kffacxfA7G4") # 拼写错误:conten → content req = r.content soup = BeautifulSoup(req,'html.parser') print(soup.prettify()) # all是Python内置关键字,换个变量名避免冲突 comment_containers = soup.find_all('div',{'id' : 'contents'})
不过就算修正了这些,你还是拿不到评论——这是因为YouTube的评论是通过JavaScript动态加载的:requests只能获取页面的静态HTML源码,而评论内容是页面加载完成后,通过AJAX请求异步拉取渲染的,所以BeautifulSoup根本看不到这些动态生成的元素。
两种可行的解决方案
方案1:用Selenium模拟浏览器加载(适合快速上手)
Selenium可以模拟真实浏览器的行为,等页面完全加载(包括动态内容)后再抓取源码,这样就能拿到评论了。
首先安装依赖:
pip install selenium
(还要下载对应浏览器的驱动,比如ChromeDriver,记得把路径配置正确)
示例代码:
from selenium import webdriver from selenium.webdriver.chrome.service import Service from bs4 import BeautifulSoup import time # 初始化Chrome浏览器,替换成你的ChromeDriver路径 service = Service('path/to/chromedriver') driver = webdriver.Chrome(service=service) # 打开目标视频页面 driver.get("https://www.youtube.com/watch?v=kffacxfA7G4") # 等待页面加载,模拟滚动加载更多评论(YouTube需要滚动才会加载更多) time.sleep(3) # 先等3秒让基础内容加载 # 滚动3次,每次加载后等2秒 for _ in range(3): driver.execute_script("window.scrollTo(0, document.documentElement.scrollHeight);") time.sleep(2) # 获取完全加载后的页面源码 page_source = driver.page_source soup = BeautifulSoup(page_source, 'html.parser') # 定位评论容器(YouTube的class可能会更新,记得检查页面结构) comments = soup.find_all('ytd-comment-thread-renderer', class_='style-scope ytd-comment-section-renderer') # 提取评论内容 for comment in comments: author = comment.find('span', class_='style-scope ytd-comment-renderer').text content = comment.find('yt-formatted-string', class_='style-scope ytd-comment-renderer').text print(f"作者:{author}\n评论:{content}\n---") # 关闭浏览器 driver.quit()
方案2:使用YouTube Data API(更稳定可靠)
如果长期爬取评论,官方API是更好的选择——不会因为页面结构变化而失效,还能直接拿到结构化的数据。
步骤大概是:
- 去Google Cloud控制台创建项目,启用YouTube Data API v3
- 获取你的API密钥
- 用
requests调用API获取评论
示例代码:
import requests API_KEY = "你的API密钥" VIDEO_ID = "kffacxfA7G4" # 调用API的接口,maxResults可以调整单次获取的评论数量 url = f"https://www.googleapis.com/youtube/v3/commentThreads?part=snippet&videoId={VIDEO_ID}&key={API_KEY}&maxResults=100" response = requests.get(url) data = response.json() # 提取评论内容 for item in data['items']: comment_info = item['snippet']['topLevelComment']['snippet'] author = comment_info['authorDisplayName'] content = comment_info['textDisplay'] print(f"作者:{author}\n评论:{content}\n---")
注意:API有免费调用额度,足够个人使用,若需要更多请求可以升级付费套餐。
内容的提问来源于stack exchange,提问作者Abhilash
相关产品推荐
相关产品推荐

