You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Colab中用Selenium爬取Youtube遇报错,求代码修改方案

解决Youtube爬取的滚动报错问题

问题原因

你代码里用By.CLASS_NAME, 'body'定位元素是错误的——Youtube页面的body标签并没有body这个类名,应该通过标签名来定位。另外,页面加载需要时间,直接执行操作会因元素未加载完成报错。

修改后的完整代码

!pip install selenium
!apt-get update
!apt install chromium-chromedriver
!cp /usr/lib/chromium-browser/chromedriver /usr/bin

from selenium import webdriver as wb
from selenium.webdriver.common.keys import Keys
from bs4 import BeautifulSoup as bs
import pandas as pd
import time
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

options = wb.ChromeOptions()
options.add_argument('--headless')     
options.add_argument('--no-sandbox')
options.add_argument('--disable-dev-shm-usage')
driver = wb.Chrome('chromedriver', options=options)

url = "https://www.youtube.com/results?search_query=%EB%94%A9%EA%B3%A0%EB%AE%A4%EC%A7%81" 
driver.get(url)

# 等待页面加载完成,通过标签名定位body元素
wait = WebDriverWait(driver, 10)
body = wait.until(EC.presence_of_element_located((By.TAG_NAME, 'body')))
body.send_keys(Keys.PAGE_DOWN)

关键修改点

  • 替换定位方式:把By.CLASS_NAME, 'body'改为By.TAG_NAME, 'body',因为body是HTML标签名而非类名
  • 添加等待机制:用WebDriverWait确保元素加载完成后再执行操作,避免因页面未渲染完成报错
  • 如果需要多次滚动加载内容,可以用循环实现:
# 示例:滚动5次,每次间隔1秒等待内容加载
for _ in range(5):
    body.send_keys(Keys.PAGE_DOWN)
    time.sleep(1)

内容的提问来源于stack exchange,提问作者dani

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 09:20:28