Selenium脚本无法点击YouTube描述的“...more”按钮并提取时间戳求助
Selenium脚本无法点击YouTube描述的“...more”按钮并提取时间戳求助
看起来你的脚本虽然提示点击成功,但实际上并没有真正展开YouTube的描述内容,导致提取不到时间戳。我帮你分析下问题根源,再给你调整后的解决方案:
问题诊断
- 等待条件不够严谨:你用了
presence_of_element_located,这个条件只确认元素存在于DOM中,但不保证它已经加载完成、可见且可点击,这会导致点击操作失效。 - Headless模式窗口尺寸问题:默认的headless窗口很小,描述区域的按钮可能不在可视范围内,即使调用
scrollIntoView也可能因为窗口尺寸限制无法正常触发交互。 - Selector可能不够精准:YouTube的元素结构偶尔会调整,原有的
tp-yt-paper-button#expand可能没有匹配到真正可交互的按钮元素。
修复后的完整代码
import sys from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import TimeoutException import re import time # Set up Selenium options = webdriver.ChromeOptions() options.add_argument('--headless=new') # 使用新版headless模式,兼容性更好 options.add_argument('--window-size=1920,1080') # 设置窗口大小,确保元素可见 options.add_argument('--no-sandbox') options.add_argument('--disable-dev-shm-usage') driver = webdriver.Chrome(options=options) def extract_timestamps(url, timeout=15): try: driver.get(url) # 先等待页面核心内容加载,替代固定sleep WebDriverWait(driver, timeout).until( EC.presence_of_element_located((By.CSS_SELECTOR, "#description")) ) # 滚动到描述区域,确保按钮进入可视范围 description_container = driver.find_element(By.CSS_SELECTOR, "#description") driver.execute_script("arguments[0].scrollIntoView({block: 'center'});", description_container) time.sleep(1) # 给页面一点滚动后的稳定时间 # 等待"显示更多"按钮可点击,使用更精准的selector show_more_button = WebDriverWait(driver, timeout).until( EC.element_to_be_clickable((By.CSS_SELECTOR, "tp-yt-paper-button#expand[aria-label='Show more']")) ) # 尝试点击,优先用原生点击,不行再用JS try: show_more_button.click() print("使用原生点击成功展开描述") except Exception: print("原生点击失败,尝试JS点击") driver.execute_script("arguments[0].click();", show_more_button) # 等待描述完全展开 WebDriverWait(driver, timeout).until( EC.presence_of_element_located((By.CSS_SELECTOR, "#description-inline-expander.ytd-video-secondary-info-renderer.style-scope.expanded")) ) # 获取完整描述文本 description_element = driver.find_element(By.CSS_SELECTOR, "#description-inline-expander") description = description_element.text print("完整描述:") print(description) # 提取时间戳,支持匹配HH:MM:SS格式 timestamps = re.findall(r'\d{1,2}:\d{2}(?::\d{2})?', description) return description, timestamps except TimeoutException: print("等待元素超时,请检查网络或页面结构是否变化") return "", [] except Exception as e: print(f"提取时间戳时出错:{e}") return "", [] finally: driver.quit() # YouTube URL youtube_url = "https://youtu.be/iTmlw3vQPSs" description, timestamp_list = extract_timestamps(youtube_url) # 输出时间戳 if timestamp_list: print("\n提取到的时间戳:") for idx, timestamp in enumerate(timestamp_list, 1): print(f"{idx}. {timestamp}") else: print("\n没有在描述中找到时间戳") # 保存数据到文件 file_name = "youtube_data.txt" with open(file_name, "w") as file: file.write(f"URL: {youtube_url}\n") file.write(f"时间戳:{', '.join(timestamp_list)}\n") print(f"\n数据已保存到 {file_name}")
关键修改点说明
- 优化Headless配置:使用新版
--headless=new并设置窗口大小,避免元素因窗口过小不可见。 - 更严谨的等待条件:改用
element_to_be_clickable确保按钮可以交互,同时等待描述展开后的状态。 - 精准定位元素:给按钮selector加上
aria-label属性,避免匹配到其他同名元素。 - 滚动到描述容器:先把整个描述区域滚动到视图中间,确保按钮完全在可视范围内。
- 优化时间戳正则:支持匹配
HH:MM:SS格式的时间戳,覆盖更多场景。
备注:内容来源于stack exchange,提问作者Harsh Tripathi - blue
相关产品推荐
相关产品推荐

