You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

跨PC语音转文字异常:耳机仅单台可用,求排查方案

语音转文字转录异常问题求助

问题现象

三台PC使用语音转文字服务时,耳机表现不一致:

  • 开发PC:耳机和扬声器均能正常完成转录
  • 另外两台PC:扬声器转录正常,但耳机无法正常转录

已在Azure AI Speech-to-Text和OpenAI实时转录服务中测试,结果一致。

技术细节

  • 所有设备使用同款耳机测试
  • 使用扬声器时,三台PC转录均正常
  • 仅切换至耳机时,两台PC出现故障

已排查项

  • 音频设备识别:Windows中可正常识别耳机
  • 测试过不同型号耳机,问题依旧

请问:是否有人遇到类似问题?能否提供额外排查步骤?是什么原因导致耳机与扬声器在转录服务中的表现差异?


相关代码

import * as SpeechSDK from "microsoft-cognitiveservices-speech-sdk";

export class AudioChatUtils {
  private static transcriber: SpeechSDK.ConversationTranscriber | null = null;

  static async startRealTimeTranscription(
    speechKey: string,
    speechRegion: string,
    selectedLanguage: string,
    onRecognized: (text: string, speakerId: string) => void,
    onError: (error: string) => void,
    onPartialResult?: (partialText: string, speakerId: string) => void
  ): Promise<() => void> {
    const speechConfig = SpeechSDK.SpeechConfig.fromSubscription(speechKey, speechRegion);
    let autoDetectSourceLanguageConfig: SpeechSDK.AutoDetectSourceLanguageConfig | null = null;

    if (selectedLanguage === "auto") {
      autoDetectSourceLanguageConfig = SpeechSDK.AutoDetectSourceLanguageConfig.fromLanguages(["en-US", "ja-JP"]);
      speechConfig.setProperty(SpeechSDK.PropertyId.SpeechServiceConnection_AutoDetectSourceLanguages, "en-US,ja-JP");
    } else {
      speechConfig.speechRecognitionLanguage = selectedLanguage;
    }

    speechConfig.setProperty(SpeechSDK.PropertyId.SpeechServiceConnection_InitialSilenceTimeoutMs, "30000");
    speechConfig.setProperty(SpeechSDK.PropertyId.SpeechServiceConnection_EndSilenceTimeoutMs, "15000");
    speechConfig.setProperty(SpeechSDK.PropertyId.SpeechServiceResponse_OutputFormatOption, SpeechSDK.OutputFormat.Detailed.toString());
    speechConfig.setProperty(SpeechSDK.PropertyId.SpeechServiceResponse_PostProcessingOption, "None");
    speechConfig.setProperty(SpeechSDK.PropertyId.SpeechServiceConnection_SpeakerIdMode, "True");
    speechConfig.setProperty(SpeechSDK.PropertyId.SpeechServiceConnection_RecoMode, "Continuous");

    const audioContext = new AudioContext();
    const mediaStream = await navigator.mediaDevices.getUserMedia({ audio: true });
    const source = audioContext.createMediaStreamSource(mediaStream);

    const gainNode = audioContext.createGain();
    gainNode.gain.value = 5;

    const highpassFilter = audioContext.createBiquadFilter();
    highpassFilter.type = "highpass";
    highpassFilter.frequency.setValueAtTime(100, audioContext.currentTime);

    const bandpassFilter = audioContext.createBiquadFilter();
    bandpassFilter.type = "bandpass";
    bandpassFilter.frequency.setValueAtTime(1000, audioContext.currentTime);
    bandpassFilter.Q.setValueAtTime(0.9, audioContext.currentTime);

    const analyser = audioContext.createAnalyser();
    analyser.fftSize = 256;

    function adjustMicGain(): void {
      const buffer = new Uint8Array(analyser.frequencyBinCount);
      analyser.getByteTimeDomainData(buffer);
      const avgVolume = buffer.reduce((a, b) => a + b, 0) / buffer.length;

      let targetGain = gainNode.gain.value;
      if (avgVolume < 50) targetGain = Math.min(gainNode.gain.value + 0.5, 10);
      else if (avgVolume > 200) targetGain = Math.max(gainNode.gain.value - 0.5, 2);

      gainNode.gain.setTargetAtTime(targetGain, audioContext.currentTime, 0.1);
      requestAnimationFrame(adjustMicGain);
    }
    adjustMicGain();

    const destination = audioContext.createMediaStreamDestination();
    source
      .connect(highpassFilter)
      .connect(bandpassFilter)
      .connect(gainNode)
      .connect(analyser)
      .connect(destination);

    // Debugging: Record and play back processed audio
    const recorder = new MediaRecorder(destination.stream);
    recorder.ondataavailable = (e) => {
      const audioBlob = e.data;
      const audioURL = URL.createObjectURL(audioBlob);
      const audio = new Audio(audioURL);
      audio.play().catch((error) => {
        console.error("Audio playback failed:", error);
      });
    };
    recorder.start();

    const audioConfig = SpeechSDK.AudioConfig.fromStreamInput(destination.stream);

    if (autoDetectSourceLanguageConfig) {
      this.transcriber = SpeechSDK.ConversationTranscriber.FromConfig(speechConfig, autoDetectSourceLanguageConfig, audioConfig);
    } else {
      this.transcriber = new SpeechSDK.ConversationTranscriber(speechConfig, audioConfig);
    }

    console.log("Real-time meeting transcription started...");

    this.transcriber.transcribed = (s, e) => {
      const speakerId = e.result.speakerId || "Unknown";
      console.log(`Recognized: [Speaker ${speakerId}] ${e.result.text}`);
      onRecognized(e.result.text, speakerId);
    };

    if (onPartialResult) {
      this.transcriber.transcribing = (s, e) => {
        const speakerId = e.result.speakerId || "Unknown";
        console.log(`Partial Result: [Speaker ${speakerId}] ${e.result.text}`);
        onPartialResult(e.result.text, speakerId);
      };
    }

    this.transcriber.canceled = (s, e) => {
      console.error("Transcription canceled: " + e.errorDetails);
      this.transcriber?.stopTranscribingAsync();
      onError(`Transcription canceled: ${e.errorDetails}`);
    };

    this.transcriber.startTranscribingAsync(
      () => console.log("Listening..."),
      (err) => {
        console.trace("Error starting transcription: " + err);
        onError(`Error starting transcription: ${err}`);
      }
    );

    return () => {
      this.transcriber?.stopTranscribingAsync(
        () => {
          console.log("Real-time transcription stopped.");
          this.transcriber = null;
        },
        (err) => onError(`Error stopping transcription: ${err}`)
      );
    };
  }
}

回答

可能的原因

  1. 音频参数不匹配:部分耳机默认采样率/位深度和转录服务要求不符,而扬声器参数刚好适配,转录服务对输入音频参数有严格要求,不匹配会直接导致识别失败。
  2. 系统音频设置问题:故障PC中耳机的输入音量过低,或被标记为“通信设备”而非“录音设备”,导致转录服务获取的音频信号强度不足。
  3. Web Audio API兼容性差异:代码中的音频预处理逻辑(滤波、增益调整)在不同PC的浏览器上支持度不同,耳机作为输入源时,预处理后的音频流不符合转录服务要求。
  4. 耳机硬件模式差异:部分耳机支持单声道/立体声切换,故障PC可能默认设置为立体声,而转录服务更适配单声道输入;或耳机拾音模式在不同系统下被自动调整,影响音频质量。

额外排查步骤

  1. 调整音频设备参数:在Windows声音设置的耳机设备“高级”选项中,将采样率和位深度改为转录服务推荐值(如16kHz、16位单声道),重新测试。
  2. 跳过音频预处理:暂时注释代码中所有滤波、增益调整逻辑,直接将原始mediaStream传入转录服务,若正常则问题出在预处理环节。
  3. 测试系统录音:用Windows录音机录制耳机麦克风音频,播放确认是否清晰,若录音异常则是系统或硬件问题,与转录服务无关。
  4. 确认浏览器设备权限:检查浏览器是否允许访问耳机麦克风,调用getUserMedia时明确指定耳机设备ID(通过navigator.mediaDevices.enumerateDevices()获取设备列表),避免自动选源错误。
  5. 更新音频驱动:给故障PC更新声卡和耳机驱动,旧驱动可能存在兼容性问题。
  6. 更换浏览器测试:用Chrome、Edge等不同浏览器测试,排除浏览器对媒体设备或Web Audio API的支持差异。

代码调整建议

若排查出是预处理问题,可尝试:

  • 简化预处理逻辑,逐步添加滤波、增益调整,定位异常环节;
  • 强制转换为单声道输入(多数转录服务优先支持单声道):
    // 创建MediaStreamSource后添加单声道转换
    const splitter = audioContext.createChannelSplitter(2);
    const merger = audioContext.createChannelMerger(2);
    source.connect(splitter);
    // 取左声道合并为单声道
    splitter.connect(merger, 0, 0);
    splitter.connect(merger, 0, 1);
    // 后续连接从merger开始
    merger.connect(highpassFilter);
    

内容的提问来源于stack exchange,提问作者Su Myat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 10:00:54