You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何向Azure Speech Service(语音转文字)发送Byte Array

解决Azure Speech Service处理字节数组语音文件的问题

问题分析

你遇到的核心问题是:使用Byte Array作为输入时,Azure Speech Service无法正确识别语音,而File示例正常运行,说明音频文件本身无问题,故障出在字节数组的传递方式上。你尝试的几种方法存在以下问题:

  • 方法1/2/3:PushStream的写入时机与识别调用不同步,且未正确处理流的结束信号;方法3循环写入时错误传入整个数组而非当前字节。
  • 方法4:调用了错误的API接口(该接口用于生成语音签名,而非语音识别)。

可行解决方案

最可靠的方式是将字节数组包装为PullStream(而非PushStream),让Speech Service自行从流中读取数据,无需手动控制推送节奏。以下是完整实现:

public async Task RecognizeFromByteArray()
{
    var speechConfig = SpeechConfig.FromSubscription(_speechKey, _speechRegion);
    // 若音频格式不是默认的WAV 16kHz 16bit单声道,需手动指定格式
    // speechConfig.SetSpeechRecognitionLanguage("zh-CN"); // 可选:指定识别语言
    // AudioStreamFormat format = AudioStreamFormat.GetWaveFormatPCM(16000, 16, 1);

    byte[] byteArray = File.ReadAllBytes(_filePath);
    using var memoryStream = new MemoryStream(byteArray);
    
    // 创建PullStream,Speech Service会自动从memoryStream读取数据
    using var pullStream = AudioInputStream.CreatePullStream(
        new PullAudioInputStreamCallback(memoryStream));
    using var audioConfig = AudioConfig.FromStreamInput(pullStream);
    using var speechRecognizer = new SpeechRecognizer(speechConfig, audioConfig);

    // 订阅识别事件(适用于长语音),短语音可直接用RecognizeOnceAsync
    speechRecognizer.Recognizing += (s, e) =>
    {
        Console.WriteLine($"识别中: {e.Result.Text}");
    };

    speechRecognizer.Recognized += (s, e) =>
    {
        if (e.Result.Reason == ResultReason.RecognizedSpeech)
        {
            Console.WriteLine($"识别结果: {e.Result.Text}");
        }
        else if (e.Result.Reason == ResultReason.NoMatch)
        {
            Console.WriteLine($"未识别到语音: {e.Result.NoMatchDetails}");
        }
    };

    speechRecognizer.Canceled += (s, e) =>
    {
        Console.WriteLine($"识别取消: {e.Reason}");
        if (e.Reason == CancellationReason.Error)
        {
            Console.WriteLine($"错误详情: {e.ErrorDetails}");
        }
        speechRecognizer.StopContinuousRecognitionAsync().Wait();
    };

    // 启动连续识别(适合长语音),短语音替换为 await speechRecognizer.RecognizeOnceAsync()
    await speechRecognizer.StartContinuousRecognitionAsync();
    
    // 等待流读取完成(可根据业务场景调整等待逻辑)
    while (memoryStream.Position < memoryStream.Length)
    {
        await Task.Delay(100);
    }
    await speechRecognizer.StopContinuousRecognitionAsync();
}

// 实现PullStream的回调类
public class PullAudioInputStreamCallback : PullAudioInputStreamCallback
{
    private readonly Stream _stream;

    public PullAudioInputStreamCallback(Stream stream)
    {
        _stream = stream;
    }

    public override int Read(byte[] buffer, uint size)
    {
        return _stream.Read(buffer, 0, (int)size);
    }

    public override void Close()
    {
        _stream.Close();
    }
}

关键注意事项

  1. 音频格式匹配:确保字节数组对应的音频格式为Speech Service支持的类型(推荐WAV 16kHz 16bit单声道),格式不同时需通过AudioStreamFormat指定具体参数。
  2. 识别模式选择:RecognizeOnceAsync仅适合15秒以内的短语音,长语音需使用StartContinuousRecognitionAsync并订阅事件。
  3. 资源管理:所有Stream和Recognizer对象需用using包裹,避免资源泄漏。

内容的提问来源于stack exchange,提问作者Christina Arvig

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 02:57:55