如何在Python3.9中将Azure TTS音频输出至虚拟线缆供Discord使用?
解决Azure TTS定向输出到虚拟音频线缆的问题
步骤1:获取虚拟音频线缆的设备ID
先运行以下代码枚举所有音频输出设备,找到虚拟线缆对应的设备ID:
import azure.cognitiveservices.speech as speechsdk # 列出所有可用的音频输出设备 output_devices = speechsdk.audio.AudioDeviceInfo.enumerate_output_devices() for idx, device in enumerate(output_devices): print(f"序号: {idx}, 设备名称: {device.name}, 设备ID: {device.id}")
在输出里找到名称包含“Virtual Cable”的设备,复制它的设备ID备用。
步骤2:修改TTS代码指定输出设备
在你的tts函数中,用获取到的设备ID创建AudioConfig,替代原来的占位符:
def tts(text): speech_config = speechsdk.SpeechConfig(subscription=os.getenv('SPEECH_KEY'), region=os.getenv('SPEECH_REGION')) speech_config.speech_synthesis_voice_name = "en-US-JaneNeural" # 替换成你复制的虚拟线缆设备ID audio_config = speechsdk.audio.AudioConfig(audio_output_device_id="你的虚拟线缆设备ID") speech_synthesizer = speechsdk.SpeechSynthesizer(speech_config=speech_config, audio_config=audio_config) speech_synthesizer.speak_text_async(text).get()
步骤3:优化原有代码的小问题
调整代码里的几个不合理之处,让运行更稳定:
load_dotenv需要加括号调用,否则无法加载.env文件里的密钥- 不要把函数定义放在
while循环内部,移到循环外避免重复定义 - 增加识别结果判断,避免空文本传入TTS
修改后的完整代码:
import os from dotenv import load_dotenv import azure.cognitiveservices.speech as speechsdk # 加载环境变量 load_dotenv() def recognize_from_microphone(): speech_config = speechsdk.SpeechConfig(subscription=os.getenv('SPEECH_KEY'), region=os.getenv('SPEECH_REGION')) speech_config.speech_recognition_language="en-US" audio_config = speechsdk.audio.AudioConfig(use_default_microphone=True) speech_recognizer = speechsdk.SpeechRecognizer(speech_config=speech_config, audio_config=audio_config) print("Speak into your microphone.") speech_recognition_result = speech_recognizer.recognize_once_async().get() # 仅当识别成功时调用TTS if speech_recognition_result.reason == speechsdk.ResultReason.RecognizedSpeech: tts(speech_recognition_result.text) elif speech_recognition_result.reason == speechsdk.ResultReason.NoMatch: print(f"No speech could be recognized: {speech_recognition_result.no_match_details}") elif speech_recognition_result.reason == speechsdk.ResultReason.Canceled: cancellation_details = speech_recognition_result.cancellation_details print(f"Speech Recognition canceled: {cancellation_details.reason}") if cancellation_details.reason == speechsdk.CancellationReason.Error: print(f"Error details: {cancellation_details.error_details}") def tts(text): speech_config = speechsdk.SpeechConfig(subscription=os.getenv('SPEECH_KEY'), region=os.getenv('SPEECH_REGION')) speech_config.speech_synthesis_voice_name = "en-US-JaneNeural" # 替换成你的虚拟线缆设备ID audio_config = speechsdk.audio.AudioConfig(audio_output_device_id="你的虚拟线缆设备ID") speech_synthesizer = speechsdk.SpeechSynthesizer(speech_config=speech_config, audio_config=audio_config) speech_synthesizer.speak_text_async(text).get() # 循环执行识别 while True: recognize_from_microphone()
这样修改后,Azure TTS的音频会单独输出到虚拟线缆,不会影响其他软件的默认音频输出设备。
内容的提问来源于stack exchange,提问作者Arishia
相关产品推荐
相关产品推荐

