Azure语音转文本无法识别PC麦克风录制的WAV音频问题
问题分析与解决方案
1. 排查PC录制WAV的格式兼容性
Azure语音转文本服务对输入音频有严格格式要求,默认支持:
- 采样率:16kHz
- 位深:16位
- 声道:单声道(mono)
PC麦克风录制的WAV通常默认是44.1kHz/立体声,这类不符合要求的格式会导致服务无法识别音频内容,进而出现无响应的情况。
- 验证方法:用Audacity等音频工具打开录制的WAV,查看格式参数。
- 修复方法:将音频转换为16kHz/16位/单声道的WAV格式后再传入脚本。
2. 完善代码的错误与状态监听
你的代码缺失**canceled事件监听**,这是排查无响应问题的核心——如果音频格式不合法或出现其他服务端错误,Azure会触发canceled事件,但当前代码未处理该事件,导致Promise永远无法resolve,脚本陷入停滞。
同时建议添加超时兜底逻辑,防止因异常情况导致无限等待。修改后的代码需补充以下内容:
// 新增canceled事件监听 speechRecognizer.canceled = (s, e) => { console.log(`CANCELED: Reason=${e.reason}`); if (e.reason === sdk.CancellationReason.Error) { console.log(`CANCELED: ErrorCode=${e.errorCode}`); console.log(`CANCELED: ErrorDetails=${e.errorDetails}`); reject(new Error(`Recognition canceled: ${e.errorDetails}`)); } speechRecognizer.stopContinuousRecognitionAsync(); }; // 新增超时兜底(5分钟,可按需调整) setTimeout(() => { if (!response) { speechRecognizer.stopContinuousRecognitionAsync(); reject(new Error("Recognition timed out")); } }, 300000);
3. 优化音频读取方式
使用fs.readFileSync读取大音频文件可能引发内存占用过高的问题,建议改用流读取方式提升稳定性:
const audioConfig = sdk.AudioConfig.fromWavFileInput(fs.createReadStream(input_file));
完整修改后的代码示例
const fs = require('fs'); const sdk = require("microsoft-cognitiveservices-speech-sdk"); async function transcribe_speech(input_file) { const speechConfig = sdk.SpeechConfig.fromSubscription(process.env.SPEECH_KEY, process.env.SPEECH_REGION); speechConfig.speechRecognitionLanguage = "he-IL"; // 改用流读取音频文件 const audioConfig = sdk.AudioConfig.fromWavFileInput(fs.createReadStream(input_file)); const speechRecognizer = new sdk.SpeechRecognizer(speechConfig, audioConfig); return new Promise((resolve, reject) => { let response = ""; console.log("In promise") // 新增canceled事件监听 speechRecognizer.canceled = (s, e) => { console.log(`CANCELED: Reason=${e.reason}`); if (e.reason === sdk.CancellationReason.Error) { console.log(`CANCELED: ErrorCode=${e.errorCode}`); console.log(`CANCELED: ErrorDetails=${e.errorDetails}`); reject(new Error(`Recognition canceled: ${e.errorDetails}`)); } speechRecognizer.stopContinuousRecognitionAsync(); }; speechRecognizer.recognized = (s, e) => { console.log("In speech recognizer") if (e.result.reason === sdk.ResultReason.RecognizedSpeech) { response += e.result.text; console.log("In recognized speech") } else if (e.result.reason === sdk.ResultReason.NoMatch) { console.log("NOMATCH: Speech could not be recognized."); } }; speechRecognizer.sessionStopped = (s, e) => { console.log("Session stopped."); speechRecognizer.stopContinuousRecognitionAsync(); resolve(response); }; speechRecognizer.startContinuousRecognitionAsync(); // 超时兜底逻辑 setTimeout(() => { if (!response) { speechRecognizer.stopContinuousRecognitionAsync(); reject(new Error("Recognition timed out after 5 minutes")); } }, 300000); }); } module.exports = { transcribe_speech } transcribe_speech("./uploads/recorded_from_pc.wav") .then((msg) => { console.log("Recognition completed. Final response:"); console.log(msg); }) .catch((err) => { console.log("Recognition error:"); console.error(err); });
内容的提问来源于stack exchange,提问作者Martin Chapman
相关产品推荐
相关产品推荐

