无CC时如何生成YouTube视频字幕(Python实现)
用Python获取无CC的YouTube视频字幕方案
首先明确:无CC的YouTube视频本身没有官方提供的字幕文件,无法像带CC的视频那样直接调用API获取现成字幕。你需要通过**语音转文字(ASR)**生成字幕,以下是无需下载完整视频的实现思路和可用库:
1. 核心思路
无需下载完整MP4,只需提取视频的音频流,再通过ASR工具将音频转成文本。
2. 可用库组合示例
步骤1:提取音频流
使用pytube直接获取视频的音频流,不用下载整个视频:
from pytube import YouTube # 替换为目标视频URL yt = YouTube("https://www.youtube.com/watch?v=example_id") audio_stream = yt.streams.filter(only_audio=True).first() # 临时保存音频(也可直接处理流,无需存本地) audio_file = audio_stream.download(filename="temp_audio.mp4")
步骤2:音频转文字
推荐两个实用的ASR库:
Whisper(OpenAI出品,准确率高):
先安装依赖:pip install openai-whisper
代码示例:import whisper # 可选模型:tiny/base/small/medium/large,越大准确率越高但速度越慢 model = whisper.load_model("base") result = model.transcribe(audio_file) # 输出完整转录文本 print(result["text"]) # 输出带时间戳的字幕格式 for segment in result["segments"]: print(f"[{segment['start']:.2f}s - {segment['end']:.2f}s] {segment['text']}")SpeechRecognition(支持多引擎,如Google Web Speech API):
安装依赖:pip install SpeechRecognition pydub(需提前安装ffmpeg处理音频格式)
代码示例:import speech_recognition as sr from pydub import AudioSegment # 将MP4音频转为WAV格式(SpeechRecognition支持的格式) audio = AudioSegment.from_file(audio_file, format="mp4") audio.export("temp_audio.wav", format="wav") r = sr.Recognizer() with sr.AudioFile("temp_audio.wav") as source: audio_data = r.record(source) try: text = r.recognize_google(audio_data) print(text) except sr.UnknownValueError: print("无法识别音频内容") except sr.RequestError as e: print(f"调用Google API失败: {e}")
3. 注意事项
- Whisper支持直接处理MP4音频,无需格式转换,使用更便捷;
- 长视频建议选用
medium或large模型提升准确率,但会占用更多系统资源; - Google Web Speech API免费但有调用限制,Whisper本地运行无需API密钥,更适合批量处理场景。
内容的提问来源于stack exchange,提问作者darkracer
相关产品推荐
相关产品推荐

