基于OpenVINO的Python项目:实现MP4视频转字幕/转录文本
Hey there! Let's walk through how to build a Python project that turns MP4 videos into transcribed text or subtitles using OpenVINO. I'll break this down step by step so it's easy to follow—we'll use Whisper (a top-tier speech recognition model) since it’s fully supported by OpenVINO and delivers great accuracy.
1. 环境搭建
First, let's get your dependencies sorted. You'll need a few packages to make this work:
pip install openvino ffmpeg-python transformers openvino-dev
openvino: Core library for optimized inferenceffmpeg-python: Extracts audio from MP4 filestransformers: Handles loading and processing the Whisper modelopenvino-dev: Includes tools to download pre-trained models from OpenVINO's Model Zoo
2. 准备OpenVINO优化的Whisper模型
Whisper comes in different sizes (tiny, base, small, medium, large)—pick one based on your accuracy/speed needs. We’ll use the small model here as a balance.
Option 1: Download pre-converted model from OpenVINO Model Zoo
Run this command to grab the optimized model directly:
omz_downloader --name whisper-small
This will save the model files (whisper-small.xml and whisper-small.bin) to your local directory.
Option 2: Convert a Hugging Face Whisper model manually
If you want to use a specific Whisper variant, convert it to OpenVINO’s IR format yourself:
from transformers import WhisperForConditionalGeneration import openvino as ov # Load the Hugging Face model model = WhisperForConditionalGeneration.from_pretrained("openai/whisper-small") # Convert to OpenVINO IR format ov_model = ov.convert_model(model) # Save the optimized model ov.save_model(ov_model, "whisper-small.xml")
3. 从MP4提取音频
ASR models work with audio, so first we need to pull the audio track from your MP4. Here’s a helper function to do that (we’ll output a 16kHz mono WAV, which Whisper requires):
import ffmpeg def extract_audio(video_path, temp_audio="temp_audio.wav"): try: # Extract mono 16kHz WAV from MP4 ffmpeg.input(video_path).output(temp_audio, ac=1, ar=16000).run(quiet=True) return temp_audio except Exception as e: print(f"Oops, audio extraction failed: {e}") return None
4. 用OpenVINO运行语音识别
Now let’s write code to load the optimized model, process the audio, and generate transcriptions or subtitles.
基础转录(输出纯文本)
import openvino as ov import numpy as np from transformers import WhisperProcessor # Initialize OpenVINO core core = ov.Core() # Load and compile the model (use "GPU" instead of "CPU" if you have an Intel GPU) model = core.read_model("whisper-small.xml") compiled_model = core.compile_model(model, "CPU") # Load Whisper processor (handles audio feature conversion and text decoding) processor = WhisperProcessor.from_pretrained("openai/whisper-small") def transcribe_audio(audio_path): # Convert audio to model-compatible features audio_features, _ = processor(audio_path, return_tensors="np", sampling_rate=16000) # Run inference with OpenVINO outputs = compiled_model([audio_features])[compiled_model.output(0)] # Decode raw output to readable text transcription = processor.decode(outputs[0], skip_special_tokens=True) return transcription
生成SRT字幕(带时间戳)
If you want timed subtitles (like SRT format), we can extend the code to capture word-level timestamps:
def generate_srt(audio_path, output_srt="output.srt"): # Use OpenVINO's Whisper pipeline for timestamp support from openvino.runtime.pieces import whisper pipe = whisper.WhisperPipeline(compiled_model, processor) # Run inference with timestamp tracking result = pipe(audio_path, return_timestamps=True) # Format timestamps to SRT's HH:MM:SS,mmm format def format_srt_time(seconds): hours = int(seconds // 3600) minutes = int((seconds % 3600) // 60) secs = int(seconds % 60) millis = int((seconds - secs) * 1000) return f"{hours:02d}:{minutes:02d}:{secs:02d},{millis:03d}" # Write SRT file with open(output_srt, "w", encoding="utf-8") as f: segment_idx = 1 for chunk in result["chunks"]: start_time = format_srt_time(chunk["timestamp"][0]) end_time = format_srt_time(chunk["timestamp"][1]) f.write(f"{segment_idx}\n") f.write(f"{start_time} --> {end_time}\n") f.write(f"{chunk['text'].strip()}\n\n") segment_idx += 1 return output_srt
5. 整合完整流程
Put it all together into a single function that handles the entire workflow:
import os def video_to_text_or_subs(video_path, text_output="transcription.txt", srt_output=None): # Step 1: Extract audio from MP4 temp_audio = extract_audio(video_path) if not temp_audio: return None # Step 2: Generate transcription transcription = transcribe_audio(temp_audio) with open(text_output, "w", encoding="utf-8") as f: f.write(transcription) print(f"Transcription saved to {text_output}") # Step 3: Generate subtitles if requested if srt_output: generate_srt(temp_audio, srt_output) print(f"SRT subtitles saved to {srt_output}") # Clean up temporary audio file os.remove(temp_audio) return transcription # Example usage video_to_text_or_subs("input.mp4", "my_transcript.txt", "my_subtitles.srt")
6. 关键注意事项
- Hardware Optimization: Swap
"CPU"with"GPU"when compiling the model if you have an Intel GPU—OpenVINO will automatically optimize inference for your hardware. - Model Size: Larger Whisper models (like medium/large) give better accuracy but run slower. Start with small/tiny for testing.
- Audio Format: The code ensures we use 16kHz mono audio, which is mandatory for Whisper. Don’t skip this step!
内容的提问来源于stack exchange,提问作者Shayekh Bin Islam

