You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在不下载的情况下获取网页中Blob嵌入视频的时长?

不完整下载获取MP4视频时长的可行技术

1. 修复ffprobe/ffmpeg的403限制,仅读取元数据

你遇到的403错误大概率是因为请求未携带浏览器级别的请求头,服务器拦截了非浏览器请求。可以通过添加User-Agent、Referer等请求头绕过限制,同时ffprobe支持仅读取元数据(不会下载完整视频):

ffprobe -v error -show_entries format=duration -of default=noprint_wrappers=1:nokey=1 \
  -headers "User-Agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" \
  -headers "Referer: https://digitalcommons.usf.edu/" \
  "https://digitalcommons.usf.edu/context/tampa_natives_show/article/1134/type/native/viewcontent"

这个命令会向服务器请求视频的元数据片段,而非完整文件,返回结果为秒级的时长数值。

2. 手动发起HTTP范围请求解析MP4元数据

MP4的核心元数据存储在moov原子中,若该原子在文件开头,仅需请求前几KB数据;若在文件末尾,可先通过HEAD请求获取文件总大小,再请求末尾1-2MB数据,从中提取时长。用Python实现的核心逻辑示例:

import requests
from struct import unpack

def get_mp4_duration(url, headers):
    # 获取文件总大小
    head_resp = requests.head(url, headers=headers)
    total_size = int(head_resp.headers['Content-Length'])
    
    # 请求末尾1MB数据(覆盖moov原子大概率所在的位置)
    range_headers = headers.copy()
    range_headers['Range'] = f'bytes={total_size-1048576}-'
    resp = requests.get(url, headers=range_headers)
    
    # 解析mvhd box提取时长(简化版逻辑,适配多数MP4结构)
    data = resp.content
    offset = 0
    while offset < len(data):
        size, box_type = unpack('>I4s', data[offset:offset+8])
        if box_type == b'mvhd':
            time_scale = unpack('>I', data[offset+20:offset+24])[0]
            duration = unpack('>I', data[offset+24:offset+28])[0]
            return duration / time_scale
        offset += size
    return None

# 调用示例
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36',
    'Referer': 'https://digitalcommons.usf.edu/'
}
duration = get_mp4_duration('https://digitalcommons.usf.edu/context/tampa_natives_show/article/1134/type/native/viewcontent', headers)
print(f"视频时长:{duration}秒")

3. 使用支持HTTP范围请求的元数据工具

mediainfo是轻量级的媒体元数据解析工具,默认会利用HTTP范围请求仅获取必要的元数据,添加请求头即可绕过403:

mediainfo --Output="General;%Duration%" \
  --http-user-agent="Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" \
  --http-referer="https://digitalcommons.usf.edu/" \
  "https://digitalcommons.usf.edu/context/tampa_natives_show/article/1134/type/native/viewcontent"

返回结果为毫秒级的时长数值,可自行转换为秒。

4. 从网页前端直接提取时长

页面使用了JW Player,大概率视频元数据已通过前端脚本加载:

  • 打开浏览器开发者工具「控制台」,执行jwplayer().getDuration(),直接获取当前播放视频的时长(单位秒)。
  • 查看网页的JavaScript代码或XHR请求,寻找JW Player的配置接口(含playlist或media字段的请求),其中通常会直接返回duration字段,无需请求视频文件。

内容的提问来源于stack exchange,提问作者Liang Zhong

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 17:12:51