如何为OpenAI Whisper ASR语音识别工具提供提示短语?
Whisper语音识别:提示短语使用及安装教程
核心问题解答
Whisper支持通过**初始提示(initial prompt)**引导模型识别特定内容,用法分为命令行和Python代码两种方式:
- 命令行方式:使用
--initial_prompt参数传入提示文本
示例:如果音频涉及医疗术语,可设置提示:whisper 你的音频文件.wav --initial_prompt "这里填写你需要的提示短语,比如特定术语、语境描述等"whisper medical_recording.wav --initial_prompt "这段音频包含冠心病、血常规等医疗专业术语" - Python代码方式:在
transcribe方法中传入initial_prompt参数import whisper # 加载模型(可替换为base/small/large等) model = whisper.load_model("base") # 传入提示短语进行转录 result = model.transcribe("你的音频文件.wav", initial_prompt="提示内容") # 输出识别结果 print(result["text"])
Ubuntu 20.04 x64 LTS下的Whisper安装与基础转录步骤(基于Nvidia GeForce RTX 3090测试)
- 创建并激活conda环境:
conda create -y --name whisperpy39 python==3.9 conda activate whisperpy39 - 安装Whisper及依赖工具:
pip install git+https://github.com/openai/whisper.git sudo apt update && sudo apt install ffmpeg - 执行转录:
# 使用默认模型转录 whisper recording.wav # 使用大模型提升识别精度 whisper recording.wav --model large
Nvidia GeForce RTX 3090显卡适配补充步骤
激活conda环境后,需安装适配CUDA的PyTorch版本以利用显卡加速:
pip install -f https://download.pytorch.org/whl/torch_stable.html conda install pytorch==1.10.1 torchvision torchaudio cudatoolkit=11.0 -c pytorch
内容的提问来源于stack exchange,提问作者Franck Dernoncourt
相关产品推荐
相关产品推荐

