You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure Communication Services语音转文本集成问题求助

解决方案:ACS电话机器人语音识别问题处理

一、解决用户回复无法捕获与超时异常问题

1. 修正StartRecognizingAsync的使用逻辑

你的代码中调用StartRecognizingAsync后直接返回模拟结果是错误的——该方法仅启动识别操作,不会等待识别完成。真实识别结果需要通过ACS的事件回调(如RecognizeCompleted事件)接收,而非同步获取返回值。

正确处理流程可借助TaskCompletionSource异步等待识别结果:

async Task<string> RecognizeSpeechAsync(CallMedia callMedia, string callerId, ILogger logger, TaskCompletionSource<string> tcs)
{
    var recognizeOptions = new CallMediaRecognizeSpeechOptions(CommunicationIdentifier.FromRawId(callerId))
    {
        InitialSilenceTimeout = TimeSpan.FromSeconds(10),
        EndSilenceTimeout = TimeSpan.FromSeconds(2), // 调整为更合理的结束沉默时长
        OperationContext = "SpeechRecognition_" + Guid.NewGuid(), // 用唯一标识关联事件
        SpeechLanguage = "zh-CN" // 明确指定语音语言,提升识别准确性
    };

    try
    {
        await callMedia.StartRecognizingAsync(recognizeOptions);
        logger.LogInformation("语音识别已启动,等待用户回复");
        return await tcs.Task; // 等待事件回调触发完成
    }
    catch (Exception ex)
    {
        logger.LogError("启动语音识别失败: {Message}", ex.Message);
        tcs.TrySetResult(string.Empty);
        return string.Empty;
    }
}

// 在事件处理方法中完成TaskCompletionSource
public void HandleRecognizeCompleted(RecognizeCompletedEventData eventData)
{
    if (eventData.OperationContext.StartsWith("SpeechRecognition_"))
    {
        var tcs = GetTaskCompletionSourceByOperationContext(eventData.OperationContext); // 自行维护OperationContext与TCS的映射关系
        if (eventData.RecognizeResult.SpeechRecognitionResult != null)
        {
            tcs.TrySetResult(eventData.RecognizeResult.SpeechRecognitionResult.DisplayText);
        }
        else if (eventData.RecognizeResult.RecognitionFailureReason == RecognitionFailureReason.InitialSilenceTimeout)
        {
            tcs.TrySetResult(string.Empty); // 标记为无回复
        }
    }
}

2. 优化超时配置与语音语言设置

  • 明确设置SpeechLanguage:根据用户使用语言指定(如zh-CN、en-US),避免默认语言不匹配导致识别失效。
  • 调整EndSilenceTimeout:5秒过长,建议设为2-3秒,减少不必要等待;InitialSilenceTimeout保持10秒即可,但需确保事件回调正确监听。

二、实现未检测到回复时的重新提示逻辑

在识别结果为空(或触发初始沉默超时)时,循环执行提示音播放+识别操作,直到获取有效回复或达到最大重试次数:

async Task<string> GetValidUserResponseAsync(CallMedia callMedia, string callerId, ILogger logger, int maxRetries = 3)
{
    int retryCount = 0;
    while (retryCount < maxRetries)
    {
        var tcs = new TaskCompletionSource<string>();
        // 播放提示语
        await callMedia.PlayAsync(new PlaySource[] { new TextSource("请说出您的回复") }, CommunicationIdentifier.FromRawId(callerId));
        // 启动识别
        var response = await RecognizeSpeechAsync(callMedia, callerId, logger, tcs);
        if (!string.IsNullOrEmpty(response))
        {
            return response;
        }
        retryCount++;
        // 播放重试提示
        await callMedia.PlayAsync(new PlaySource[] { new TextSource("未听到您的回复,请再试一次") }, CommunicationIdentifier.FromRawId(callerId));
    }
    logger.LogWarning("用户多次未回复,终止流程");
    return string.Empty;
}

三、提升语音转文本准确性的建议

  • 指定精准语音语言:根据目标用户群体设置SpeechLanguage,避免自动检测的误差。
  • 使用自定义语音模型:如果业务涉及专业术语,在Azure Speech Studio训练自定义语音模型,然后在CallMediaRecognizeSpeechOptions中指定SpeechModelId。
  • 优化音频质量:确保ACS电话线路音频稳定,避免背景噪音;Speech Services默认集成噪音抑制功能,确认部署区域支持即可。
  • 启用连续识别(按需):若用户需长时间发言,开启CallMediaRecognizeSpeechOptions的EnableContinuousRecognition属性(电话场景通常用单次识别即可)。

四、ACS语音识别的替代实现方案

如果默认集成的Speech Services无法满足需求,可考虑两种方向:

  • 直接调用Azure Speech Services SDK处理音频流:从ACS获取呼叫音频流,传入Speech Services的SpeechRecognizer进行识别,这种方式更灵活,可自定义脏话过滤、实时中间结果等参数。
  • Bot Framework + ACS集成:借助Bot Framework成熟的对话管理能力处理语音交互流程,降低自定义逻辑的复杂度。

内容的提问来源于stack exchange,提问作者Saad Illahi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 02:14:55