使用gTTS实现多语种对话语音合成报错求助
问题排查与解决
核心错误原因
这个FileNotFoundError是因为Pydub依赖FFmpeg工具处理音频文件,但你的系统里没有安装FFmpeg,或者FFmpeg的路径没添加到系统环境变量中。Pydub本身不具备音频编解码能力,它需要调用FFmpeg的ffmpeg和ffprobe命令来读取、处理音频,系统找不到这些命令就会触发这个错误。
另外你的代码还有个逻辑bug:循环中Person1的内容每次都保存为temp1.mp3,后续的Person1内容会覆盖之前的文件,导致合并后的音频里重复出现最后一段Person1的内容,而不是正确的对话顺序。
解决步骤
1. 安装并配置FFmpeg
- 下载Windows版FFmpeg静态编译包(去FFmpeg官方下载对应你系统位数的版本)
- 解压压缩包,找到里面的
bin文件夹(里面包含ffmpeg.exe、ffprobe.exe等文件) - 把这个
bin文件夹的路径添加到系统环境变量的Path中 - 重启你的命令行或Python IDE,让环境变量生效
2. 修改代码修复临时文件覆盖问题
把固定的临时文件名改成动态生成的唯一名称,避免覆盖。下面是修改后的完整代码:
import os from gtts import gTTS from pydub import AudioSegment # Dialog between two people dialog = "Person 1: Hi, how are you? \nPerson 2: I'm good, thanks for asking. How about you? \nPerson 1: I'm great, thanks!" # Split the dialog into separate lines lines = dialog.split("\n") # Create an empty list to store the mp3 files mp3_files = [] # Convert each line of dialog to speech and save as mp3 for idx, line in enumerate(lines): temp_file = f"temp_{idx}.mp3" if "Person 1" in line: # Use English voice for Person 1 tts = gTTS(text=line, lang='en', slow=False) tts.save(temp_file) mp3_files.append(temp_file) elif "Person 2" in line: # Use French voice for Person 2 tts = gTTS(text=line, lang='fr', slow=False) tts.save(temp_file) mp3_files.append(temp_file) # Create an empty audio segment combined = AudioSegment.empty() # Iterate through the list of mp3 files and concatenate them for file in mp3_files: combined += AudioSegment.from_file(file, format="mp3") # Export the combined audio segment to a new mp3 file combined.export("dialog.mp3", format="mp3") # Remove the temp files for file in mp3_files: os.remove(file)
代码修改说明
- 用
enumerate获取循环索引,生成temp_0.mp3、temp_1.mp3这类唯一的临时文件名,避免内容覆盖 - 移除了没用的
moviepy导入(你代码里没用到moviepy的功能,留着没用)
验证
完成上述步骤后,重新运行代码,就能正常生成合并后的dialog.mp3文件,且对话内容顺序正确,没有重复。
内容的提问来源于stack exchange,提问作者Electro
相关产品推荐
相关产品推荐

