You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于OpenVINO的Python项目:实现MP4视频转字幕/转录文本

使用OpenVINO实现MP4视频转文本/字幕的Python方案

Hey there! Let's walk through how to build a Python project that turns MP4 videos into transcribed text or subtitles using OpenVINO. I'll break this down step by step so it's easy to follow—we'll use Whisper (a top-tier speech recognition model) since it’s fully supported by OpenVINO and delivers great accuracy.

1. 环境搭建

First, let's get your dependencies sorted. You'll need a few packages to make this work:

pip install openvino ffmpeg-python transformers openvino-dev
  • openvino: Core library for optimized inference
  • ffmpeg-python: Extracts audio from MP4 files
  • transformers: Handles loading and processing the Whisper model
  • openvino-dev: Includes tools to download pre-trained models from OpenVINO's Model Zoo

2. 准备OpenVINO优化的Whisper模型

Whisper comes in different sizes (tiny, base, small, medium, large)—pick one based on your accuracy/speed needs. We’ll use the small model here as a balance.

Option 1: Download pre-converted model from OpenVINO Model Zoo

Run this command to grab the optimized model directly:

omz_downloader --name whisper-small

This will save the model files (whisper-small.xml and whisper-small.bin) to your local directory.

Option 2: Convert a Hugging Face Whisper model manually

If you want to use a specific Whisper variant, convert it to OpenVINO’s IR format yourself:

from transformers import WhisperForConditionalGeneration
import openvino as ov

# Load the Hugging Face model
model = WhisperForConditionalGeneration.from_pretrained("openai/whisper-small")
# Convert to OpenVINO IR format
ov_model = ov.convert_model(model)
# Save the optimized model
ov.save_model(ov_model, "whisper-small.xml")

3. 从MP4提取音频

ASR models work with audio, so first we need to pull the audio track from your MP4. Here’s a helper function to do that (we’ll output a 16kHz mono WAV, which Whisper requires):

import ffmpeg

def extract_audio(video_path, temp_audio="temp_audio.wav"):
    try:
        # Extract mono 16kHz WAV from MP4
        ffmpeg.input(video_path).output(temp_audio, ac=1, ar=16000).run(quiet=True)
        return temp_audio
    except Exception as e:
        print(f"Oops, audio extraction failed: {e}")
        return None

4. 用OpenVINO运行语音识别

Now let’s write code to load the optimized model, process the audio, and generate transcriptions or subtitles.

基础转录(输出纯文本)

import openvino as ov
import numpy as np
from transformers import WhisperProcessor

# Initialize OpenVINO core
core = ov.Core()
# Load and compile the model (use "GPU" instead of "CPU" if you have an Intel GPU)
model = core.read_model("whisper-small.xml")
compiled_model = core.compile_model(model, "CPU")

# Load Whisper processor (handles audio feature conversion and text decoding)
processor = WhisperProcessor.from_pretrained("openai/whisper-small")

def transcribe_audio(audio_path):
    # Convert audio to model-compatible features
    audio_features, _ = processor(audio_path, return_tensors="np", sampling_rate=16000)
    # Run inference with OpenVINO
    outputs = compiled_model([audio_features])[compiled_model.output(0)]
    # Decode raw output to readable text
    transcription = processor.decode(outputs[0], skip_special_tokens=True)
    return transcription

生成SRT字幕(带时间戳)

If you want timed subtitles (like SRT format), we can extend the code to capture word-level timestamps:

def generate_srt(audio_path, output_srt="output.srt"):
    # Use OpenVINO's Whisper pipeline for timestamp support
    from openvino.runtime.pieces import whisper
    pipe = whisper.WhisperPipeline(compiled_model, processor)
    # Run inference with timestamp tracking
    result = pipe(audio_path, return_timestamps=True)

    # Format timestamps to SRT's HH:MM:SS,mmm format
    def format_srt_time(seconds):
        hours = int(seconds // 3600)
        minutes = int((seconds % 3600) // 60)
        secs = int(seconds % 60)
        millis = int((seconds - secs) * 1000)
        return f"{hours:02d}:{minutes:02d}:{secs:02d},{millis:03d}"

    # Write SRT file
    with open(output_srt, "w", encoding="utf-8") as f:
        segment_idx = 1
        for chunk in result["chunks"]:
            start_time = format_srt_time(chunk["timestamp"][0])
            end_time = format_srt_time(chunk["timestamp"][1])
            f.write(f"{segment_idx}\n")
            f.write(f"{start_time} --> {end_time}\n")
            f.write(f"{chunk['text'].strip()}\n\n")
            segment_idx += 1
    return output_srt

5. 整合完整流程

Put it all together into a single function that handles the entire workflow:

import os

def video_to_text_or_subs(video_path, text_output="transcription.txt", srt_output=None):
    # Step 1: Extract audio from MP4
    temp_audio = extract_audio(video_path)
    if not temp_audio:
        return None

    # Step 2: Generate transcription
    transcription = transcribe_audio(temp_audio)
    with open(text_output, "w", encoding="utf-8") as f:
        f.write(transcription)
    print(f"Transcription saved to {text_output}")

    # Step 3: Generate subtitles if requested
    if srt_output:
        generate_srt(temp_audio, srt_output)
        print(f"SRT subtitles saved to {srt_output}")

    # Clean up temporary audio file
    os.remove(temp_audio)
    return transcription

# Example usage
video_to_text_or_subs("input.mp4", "my_transcript.txt", "my_subtitles.srt")

6. 关键注意事项

  • Hardware Optimization: Swap "CPU" with "GPU" when compiling the model if you have an Intel GPU—OpenVINO will automatically optimize inference for your hardware.
  • Model Size: Larger Whisper models (like medium/large) give better accuracy but run slower. Start with small/tiny for testing.
  • Audio Format: The code ensures we use 16kHz mono audio, which is mandatory for Whisper. Don’t skip this step!

内容的提问来源于stack exchange,提问作者Shayekh Bin Islam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 09:16:21