You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

JS MediaRecorder音频流对接Azure语音转写无结果问题排查

问题:MediaRecorder流式音频无法被Azure实时语音转写API识别

我有两段JavaScript客户端代码,通过WebSocket将麦克风音频流式传输到Python服务器,再转发给Azure实时语音转写API完成转写与说话人分离:

  • 基于ScriptProcessor的实现能正常生成转写结果,但CPU占用过高
  • 切换到MediaRecorder后,Azure始终返回无结果,无法完成转写

两段代码核心差异:ScriptProcessor按字节大小切割原始PCM音频,MediaRecorder按时长切割封装后的压缩音频(默认WebM/Opus格式)


核心原因

Azure实时语音转写API仅支持无封装的原始PCM音频流,要求格式为:16kHz采样率、单声道、16位有符号整数、小端字节序。而MediaRecorder默认输出的是经过封装压缩的音频数据,服务器直接将这类数据推送给Azure时,API无法解析,因此返回无结果。


解决方案

一、客户端修改:让MediaRecorder输出原始PCM格式

修改MediaRecorder配置,指定输出带WAV封装的PCM音频(WAV头包含格式信息,方便服务器解析):

const connectButton = document.getElementById("connectButton");
const startButton = document.getElementById("startButton");
const stopButton = document.getElementById("stopButton");
let mediaRecorder;
let socket;

connectButton.addEventListener("click", () => {
  socket = new WebSocket("ws://localhost:8000");

  socket.addEventListener("open", () => {
    console.log("Connected to server");
    connectButton.disabled = true;
    startButton.disabled = false;
  });

  socket.addEventListener("close", () => {
    console.log("Disconnected from server");
    connectButton.disabled = false;
    startButton.disabled = true;
    stopButton.disabled = true;
  });
});

startButton.addEventListener("click", async () => {
  const stream = await navigator.mediaDevices.getUserMedia({ audio: true });
  // 指定输出WAV封装的PCM音频
  const pcmOptions = { mimeType: 'audio/wav;codec=pcm' };
  
  // 兼容性校验:浏览器不支持则回退到默认格式
  if (!MediaRecorder.isTypeSupported(pcmOptions.mimeType)) {
    console.warn(`${pcmOptions.mimeType} not supported, falling back to default`);
    mediaRecorder = new MediaRecorder(stream);
  } else {
    mediaRecorder = new MediaRecorder(stream, pcmOptions);
  }

  mediaRecorder.ondataavailable = (event) => {
    if (event.data.size > 0 && socket && socket.readyState === WebSocket.OPEN) {
      socket.send(event.data);
      console.log("audio chunk sent");
    }
  };

  mediaRecorder.start(100); // 按100ms分片传输

  startButton.disabled = true;
  stopButton.disabled = false;
});

stopButton.addEventListener("click", () => {
  if (mediaRecorder) {
    mediaRecorder.stop();
  }
  if (socket) {
    socket.close();
  }
  startButton.disabled = false;
  stopButton.disabled = true;
});

二、服务器修改:解析WAV头并适配Azure格式要求

服务器需要先从WAV数据中提取纯PCM,再转换为Azure要求的格式:

1. 新增WAV头解析函数

def extract_pcm_from_wav(wav_bytes):
    """从WAV字节数据中提取纯PCM音频,跳过44字节的WAV头"""
    if len(wav_bytes) < 44:
        return b""
    
    # 解析WAV头中的关键参数(可选,用于校验格式)
    sample_rate = int.from_bytes(wav_bytes[24:28], byteorder='little')
    num_channels = int.from_bytes(wav_bytes[22:24], byteorder='little')
    bits_per_sample = int.from_bytes(wav_bytes[34:36], byteorder='little')
    
    print(f"Received WAV: sample_rate={sample_rate}, channels={num_channels}, bits={bits_per_sample}")
    
    # 返回纯PCM数据
    return wav_bytes[44:]

2. 修改客户端连接处理逻辑

更新handle_client_connection函数,添加WAV解析、格式转换逻辑:

async def handle_client_connection(websocket, path):
    global write_stream
    global buffer
    global write_stream_sampled

    print("Client connected")
    transcriber, push_stream = setup_azure_service()
    transcriber.start_transcribing_async().get()

    try:
        async for message in websocket:
            if buffer is None:
                buffer = b""
            
            if write_stream is None:
                    write_stream = open("output.wav", "ab")
            
            if write_stream_sampled is None:
                write_stream_sampled = open("output_sampled.pcm", "ab")
            
            if isinstance(message, bytes):
                write_stream.write(message)
                # 提取纯PCM数据
                pcm_data = extract_pcm_from_wav(message)
                if not pcm_data:
                    continue
                buffer += pcm_data
                
                # 将PCM转换为Azure要求的16kHz单声道格式
                # 假设原始采样率为44100,根据实际情况调整
                downsampled_pcm = downsample_audio(buffer, 44100, 16000, num_channels=1)
                # 推送给Azure转写API
                push_stream.write(downsampled_pcm)
                buffer = b""  # 清空缓冲区避免堆积
                print(f"Processed PCM chunk size: {len(downsampled_pcm)}")
    except websockets.ConnectionClosed:
        print("Client disconnected")
    finally:
        if write_stream:
            write_stream.close()
            write_stream = None
        
        transcriber.stop_transcribing_async().get()

3. 优化降采样函数(确保输出符合要求)

def downsample_audio(byte_chunk, original_rate, target_rate, num_channels=1):
    audio_data = np.frombuffer(byte_chunk, dtype=np.int16)
    
    if num_channels == 2:
        # 立体声转单声道:取左右声道平均值
        audio_data = audio_data.reshape(-1, 2)
        audio_data = np.mean(audio_data, axis=1).astype(np.int16)
    
    num_samples = int(len(audio_data) * target_rate / original_rate)
    downsampled_audio = resample(audio_data, num_samples)
    downsampled_audio = np.round(downsampled_audio).astype(np.int16)
    return downsampled_audio.tobytes()

验证步骤

  1. 启动修改后的Python服务器
  2. 运行MediaRecorder客户端,点击「连接」和「开始录制」
  3. 查看服务器日志,确认输出Received WAV: sample_rate=xxx, channels=xxx
  4. 说话测试,服务器应输出类似以下的转写结果:
Session started
Language: en-US
    Text=Hello, this is a test
    Speaker ID=1
Language: en-US
    Text=Azure speech to text works with MediaRecorder now
    Speaker ID=1

关键注意事项

  • 浏览器兼容性:若浏览器不支持audio/wav;codec=pcm,需额外用FFmpeg等工具在服务器端解码压缩音频(如Opus)为PCM
  • 格式一致性:必须确保推送给Azure的是16kHz采样率、单声道、16位有符号整数、小端字节序的PCM数据
  • 缓冲区管理:处理完音频数据后及时清空缓冲区,避免内存堆积

内容的提问来源于stack exchange,提问作者Googler Thiru

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.18 00:15:54