You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure Speech to Text自动语言检测功能无法正常工作求助

问题:NextJS中Azure Speech服务自动语言检测的连续语音转文字失效

基于NextJS开发项目,已成功部署单语言Speech-to-Text功能,但带自动语言检测的连续语音转文字功能始终无法正常运行。参考官方文档配置了AutoDetectSourceLanguageConfig,但问题依旧,且React/JavaScript相关的示例资料较少。

用户代码片段:

useEffect(() => {
  const fetchTokenAndSetupRecognizer = async () => {
    const tokenObj = await getTokenOrRefresh();
    if (tokenObj.authToken && tokenObj.region) {
      audioConfig.current = AudioConfig.fromDefaultMicrophoneInput();

      const autoDetectLanguages = [
        "en-US",
        "de-DE"
      ];
      speechConfig.current = SpeechConfig.fromAuthorizationToken(
        tokenObj.authToken,
        tokenObj.region
      );
      const autoDetectConfig =
        AutoDetectSourceLanguageConfig.fromLanguages(autoDetectLanguages);

      audioConfig.current = AudioConfig.fromDefaultMicrophoneInput();
      recognizer.current = SpeechRecognizer.FromConfig(
        speechConfig.current,
        autoDetectConfig,
        audioConfig.current
      );
      recognizer.current.recognized = (s, e) =>
        processRecognizedTranscript(e);
      recognizer.current.canceled = (s, e) => handleCanceled(e);
    }
    setIsDisabled(!recognizer.current);
  };
  fetchTokenAndSetupRecognizer();
  return () => {
    recognizer.current?.close();
  };
}, []);

排查与解决方案

  • 确认Speech SDK版本
    确保使用最新版azure-cognitiveservices-speech包,旧版本可能存在兼容性问题。执行以下命令检查并升级:

    npm list azure-cognitiveservices-speech
    npm install azure-cognitiveservices-speech@latest
    
  • 启动连续识别流程
    代码中仅初始化了识别器,但未启动连续识别。需在配置完成后调用启动方法:

    // 初始化recognizer后添加
    recognizer.current.startContinuousRecognitionAsync(
      () => console.log("连续识别已启动"),
      (err) => console.error("启动失败:", err)
    );
    
  • 正确获取语言检测结果
    自动语言检测的结果需从识别结果的properties中提取,修改processRecognizedTranscript函数:

    import { PropertyId } from "azure-cognitiveservices-speech";
    
    const processRecognizedTranscript = (e) => {
      if (e.result.reason === ResultReason.RecognizedSpeech) {
        // 获取检测到的语言代码
        const detectedLang = e.result.properties.getProperty(PropertyId.SpeechServiceConnection_AutoDetectSourceLanguageResult);
        console.log("检测到语言:", detectedLang);
        console.log("识别文本:", e.result.text);
        // 后续业务逻辑处理
      }
    };
    
  • 优化SpeechConfig配置
    显式设置连续识别相关参数,避免默认配置冲突:

    speechConfig.current.setSpeechRecognitionOutputFormat(OutputFormat.Detailed);
    // 设置端点静音超时(根据需求调整,单位:毫秒)
    speechConfig.current.setProperty(PropertyId.SpeechServiceConnection_EndSilenceTimeoutMs, "5000");
    
  • 完善组件清理逻辑
    在组件卸载前先停止识别再关闭识别器,避免资源泄漏:

    return () => {
      if (recognizer.current) {
        recognizer.current.stopContinuousRecognitionAsync();
        recognizer.current.close();
      }
    };
    
  • 验证资源权限与区域
    确认Azure Speech资源的区域支持自动语言检测功能,且token对应的权限包含语音转文字和语言识别权限,可通过Azure门户检查资源状态。


内容的提问来源于stack exchange,提问作者jojak

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.18 05:22:05