You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python的Selenium抓取任意URL视频遇到问题求助

通用网页视频抓取问题解决方案

问题分析

  • 静态爬虫(BeautifulSoup)局限:只能获取HTML静态内容,现代网站视频多通过JS动态渲染、流媒体协议加载,静态请求无法拿到真实视频地址。你原代码里查找video标签内a标签href的逻辑错误——视频地址通常存于video标签的src属性,或内部source标签的src属性,而非嵌套的a标签。
  • Selenium代码问题:
    • 使用find_element(单数方法)会在无匹配元素时直接抛出异常,应使用find_elements返回空列表而非报错。
    • 页面未完全加载就执行元素查找,缺少等待逻辑。
    • 原生无头模式易被网站检测,导致内容加载异常。

改进后的代码示例

1. 优化BeautifulSoup代码(仅适用于静态嵌入视频)

import requests
from bs4 import BeautifulSoup

Web_url = "目标URL"
try:
    # 添加User-Agent模拟浏览器请求
    headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36"}
    r = requests.get(Web_url, headers=headers)
    r.raise_for_status()  # 捕获请求错误
    
    soup = BeautifulSoup(r.content, 'html.parser')
    # 查找所有video标签
    video_tags = soup.find_all('video')
    
    for video in video_tags:
        # 提取video标签自身的src
        video_src = video.get('src')
        if video_src:
            print(f"静态视频地址: {video_src}")
        # 提取内部source标签的src
        source_tags = video.find_all('source')
        for source in source_tags:
            source_src = source.get('src')
            if source_src:
                print(f"Source标签视频地址: {source_src}")
except Exception as e:
    print(f"请求或解析错误: {e}")

2. 优化Selenium代码(处理动态加载视频)

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

url = "目标URL"
chrome_options = Options()
# 优化无头模式,避免被网站检测
chrome_options.add_argument("--headless=new")
chrome_options.add_argument("--disable-blink-features=AutomationControlled")
chrome_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36")

driver = webdriver.Chrome(options=chrome_options)
try:
    driver.get(url)
    # 等待页面加载,最多等待10秒
    WebDriverWait(driver, 10).until(EC.presence_of_all_elements_located((By.TAG_NAME, "video")))
    
    # 获取所有video元素
    videos = driver.find_elements(By.TAG_NAME, "video")
    if not videos:
        print("未找到video元素,尝试查找iframe内的视频")
        # 遍历所有iframe,切换后查找视频
        iframes = driver.find_elements(By.TAG_NAME, "iframe")
        for iframe in iframes:
            driver.switch_to.frame(iframe)
            iframe_videos = driver.find_elements(By.TAG_NAME, "video")
            for vid in iframe_videos:
                src = vid.get_attribute('src')
                if src:
                    print(f"iframe内视频地址: {src}")
            driver.switch_to.default_content()
    else:
        for video in videos:
            src = video.get_attribute('src')
            if src:
                print(f"视频地址: {src}")
            # 提取video内部source标签的地址
            sources = video.find_elements(By.TAG_NAME, "source")
            for source in sources:
                source_src = source.get_attribute('src')
                if source_src:
                    print(f"Source视频地址: {source_src}")
except Exception as e:
    print(f"错误信息: {e}")
finally:
    driver.quit()

通用抓取思路

  • 动态内容处理:多数网站视频地址通过JS动态生成,需等待页面完全渲染,必要时模拟滚动触发加载。
  • 流媒体协议识别:很多视频采用HLS(.m3u8)或DASH协议,可通过浏览器开发者工具的「网络」面板筛选media类型,查看真实播放列表请求。
  • 反爬应对:添加合理的User-Agent,优化无头模式参数,必要时配合代理(通用场景需根据网站调整)。
  • 局限性说明:不存在万能的通用视频抓取方案,不同网站的视频加载逻辑差异极大,针对特定网站需分析其接口请求逻辑(如XHR接口返回的视频地址)。

内容的提问来源于stack exchange,提问作者Thomas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 22:50:33