在Colab中用Selenium爬取Youtube遇报错,求代码修改方案
解决Youtube爬取的滚动报错问题
问题原因
你代码里用By.CLASS_NAME, 'body'定位元素是错误的——Youtube页面的body标签并没有body这个类名,应该通过标签名来定位。另外,页面加载需要时间,直接执行操作会因元素未加载完成报错。
修改后的完整代码
!pip install selenium !apt-get update !apt install chromium-chromedriver !cp /usr/lib/chromium-browser/chromedriver /usr/bin from selenium import webdriver as wb from selenium.webdriver.common.keys import Keys from bs4 import BeautifulSoup as bs import pandas as pd import time from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC options = wb.ChromeOptions() options.add_argument('--headless') options.add_argument('--no-sandbox') options.add_argument('--disable-dev-shm-usage') driver = wb.Chrome('chromedriver', options=options) url = "https://www.youtube.com/results?search_query=%EB%94%A9%EA%B3%A0%EB%AE%A4%EC%A7%81" driver.get(url) # 等待页面加载完成,通过标签名定位body元素 wait = WebDriverWait(driver, 10) body = wait.until(EC.presence_of_element_located((By.TAG_NAME, 'body'))) body.send_keys(Keys.PAGE_DOWN)
关键修改点
- 替换定位方式:把
By.CLASS_NAME, 'body'改为By.TAG_NAME, 'body',因为body是HTML标签名而非类名 - 添加等待机制:用
WebDriverWait确保元素加载完成后再执行操作,避免因页面未渲染完成报错 - 如果需要多次滚动加载内容,可以用循环实现:
# 示例:滚动5次,每次间隔1秒等待内容加载 for _ in range(5): body.send_keys(Keys.PAGE_DOWN) time.sleep(1)
内容的提问来源于stack exchange,提问作者dani
相关产品推荐
相关产品推荐

