如何保存pyttsx3中runAndWait()生成的语音?求优质离线TTS库推荐
Hey there! Let's start with fixing your pyttsx3 save issue first—turns out it does support saving audio files, you just need to use the right method. Then I'll share some offline TTS libraries with way better audio quality than eSpeak.
1. 用pyttsx3保存语音到文件
You were missing the save_to_file() method in your code. Here's how to modify your function to both speak the text and save it as a WAV file:
import pyttsx3 def tts_and_save(text, output_file="output.wav"): engine = pyttsx3.init() # 可选:调整语音参数(比如语速、音量、语音包) engine.setProperty('rate', 150) # 语速,默认200 engine.setProperty('volume', 0.9) # 音量范围0-1 # 保存到文件 engine.save_to_file(text, output_file) # 执行播放与保存操作并等待完成 engine.runAndWait() # 使用示例 tts_and_save("Hello, this is saved audio!", "my_voice.wav")
注意事项:
- Windows系统需确保
pywin32已安装(pyttsx3依赖它),未安装可运行pip install pywin32。 - Linux/macOS上可能需要额外安装驱动(比如espeak),但保存功能依然可用,只是音质可能仍不如其他专业库。
2. 音质更好的离线TTS库推荐
If you're not happy with eSpeak's quality, these offline options are way better:
Coqui TTS
This is my top pick for open-source, high-quality offline TTS. It supports multiple languages, has pre-trained models that sound natural, and even lets you fine-tune your own models if needed.
基本使用步骤:
- 安装库:
pip install TTS
- 简单的文本转语音并保存示例:
from TTS.api import TTS # 选择一个预训练模型(比如英文的tts_models/en/ljspeech/tacotron2-DDC_ph) tts = TTS(model_name="tts_models/en/ljspeech/tacotron2-DDC_ph", progress_bar=False, gpu=False) # 生成并保存音频 tts.tts_to_file(text="This is a much more natural-sounding offline voice.", file_path="coqui_output.wav")
- 你可以在Coqui的模型库中找到适配不同语言的模型,很多支持中文、西班牙语等。
- 如果有GPU,开启
gpu=True会大幅加快生成速度。
Windows SAPI5 语音包(适合Windows用户)
If you're on Windows, pyttsx3 actually uses SAPI5 under the hood. You can install higher-quality offline voice packs from Microsoft (like "Microsoft David Desktop" or "Microsoft Zira Desktop") and switch to them in pyttsx3:
import pyttsx3 engine = pyttsx3.init() # 列出所有可用语音 voices = engine.getProperty('voices') for voice in voices: print(f"Voice: {voice.name}, ID: {voice.id}") # 切换到高质量语音 engine.setProperty('voice', voices[1].id) # 替换成你找到的优质语音ID engine.save_to_file("This uses a better Windows voice.", "windows_voice.wav") engine.runAndWait()
These voices are way clearer than the default eSpeak ones and work offline once downloaded.
Festival(适合Linux用户)
Festival is a classic offline TTS system for Linux, with better quality than eSpeak. You can install it via your package manager (e.g., sudo apt install festival festvox-kallpc16k for English) and use it with Python via the festival package:
pip install festival
from festival import Festival Festival().say("This is a Festival voice output.", filename="festival_output.wav")
内容的提问来源于stack exchange,提问作者zguesmi

