You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何配置Azure Speech SDK的AudioConfig仅捕获系统音频并排除麦克风输入?

如何配置Azure Speech SDK的AudioConfig仅捕获系统音频并排除麦克风输入?

嘿,我仔细看了你的问题,发现几个关键的小错误,咱们一步步把它解决掉:

1. 先纠正一个核心误解:你用错了AudioConfig的方法!

你代码里的AudioConfig.FromDefaultSpeakerOutput()是用来指定音频输出设备(比如播放识别结果的扬声器),完全不是用来捕获音频输入的!这就是为什么你的Speech SDK还在抓麦克风——因为你根本没正确指定输入源,SDK默认会用系统默认麦克风作为输入。

2. 正确思路:把JS获取的系统音频流传给Speech SDK

要只捕获屏幕共享的系统音频,你需要把JS中getDisplayMedia拿到的系统音频轨道,传递给Blazor,再让Speech SDK用这个轨道作为唯一输入源。下面是具体的实现步骤:

步骤一:修改JavaScript代码,保留系统音频并传递引用

首先,你之前的track.enabled = false会把系统音频轨道也禁用,这是完全错误的!咱们修改JS函数,提取系统音频轨道并保存,方便后续传递给Blazor:

window.startScreenSharing = async function () {
    try {
        const container = document.getElementById('sharedScreen');
        if (!container) {
            console.error('Element with ID "sharedScreen" not found');
            return null;
        }

        // 启动屏幕共享,请求视频+系统音频
        const stream = await navigator.mediaDevices.getDisplayMedia({
            video: true,
            audio: true  // 这里的audio就是系统音频,不是麦克风
        });

        // 提取系统音频轨道(getDisplayMedia返回的音频轨道就是系统声音)
        const systemAudioTrack = stream.getAudioTracks()[0];
        if (!systemAudioTrack) {
            console.error('未找到系统音频轨道');
            return null;
        }

        // 创建只包含系统音频的MediaStream,方便后续处理
        const systemAudioStream = new MediaStream([systemAudioTrack]);

        // 渲染共享屏幕的视频元素
        const videoElement = document.createElement('video');
        videoElement.srcObject = stream;
        videoElement.autoplay = true;
        videoElement.muted = true;  // 本地静音避免反馈
        videoElement.controls = false;
        videoElement.style.width = '100%';
        videoElement.style.height = '100%';
        container.innerHTML = '';
        container.appendChild(videoElement);

        // 保存全局引用,方便后续调用
        window.currentScreenStream = stream;
        window.currentSystemAudioStream = systemAudioStream;

        // 返回音频轨道ID,让Blazor可以关联到这个流
        return systemAudioTrack.id;
    } catch (err) {
        console.error('屏幕共享失败:', err);
        alert('屏幕共享失败,请检查浏览器设置');
        return null;
    }
};

步骤二:选择适合的方案对接Azure Speech SDK

这里给你两个方案,推荐用方案A,因为更适配浏览器环境:

方案A:直接用Azure Speech Web SDK(推荐)

浏览器环境下,Azure Speech Web SDK原生支持MediaStreamTrack,不用做复杂的流转换,代码更简洁:

  1. 先在index.html中引入Web SDK:
<script src="https://cdn.jsdelivr.net/npm/microsoft-cognitiveservices-speech-sdk@latest/distrib/browser/microsoft.cognitiveservices.speech.sdk.min.js"></script>
  1. 添加JS的识别控制函数:
// 启动系统音频识别
window.startSystemAudioRecognition = async function (subscriptionKey, region) {
    // 初始化Speech配置
    const speechConfig = SpeechSDK.SpeechConfig.fromSubscription(subscriptionKey, region);
    // 使用之前保存的系统音频流作为输入
    const audioConfig = SpeechSDK.AudioConfig.fromStreamInput(window.currentSystemAudioStream);
    const recognizer = new SpeechSDK.SpeechRecognizer(speechConfig, audioConfig);

    // 监听识别结果,传给Blazor
    recognizer.recognizing = (s, e) => {
        if (e.result.text.trim()) {
            window.blazorReceiveRecognitionResult(e.result.text);
        }
    };

    // 启动持续识别
    await recognizer.startContinuousRecognitionAsync();
    window.currentRecognizer = recognizer;
};

// 停止识别
window.stopSystemAudioRecognition = async function () {
    if (window.currentRecognizer) {
        await window.currentRecognizer.stopContinuousRecognitionAsync();
        window.currentRecognizer = null;
    }
};
  1. 在Blazor中调用并接收结果:
private bool isScreenSharing;
private string subscriptionKey = "你的订阅密钥";
private string region = "你的区域";

protected override async Task OnInitializedAsync()
{
    // 注册接收识别结果的回调
    await JSRuntime.InvokeVoidAsync("eval", @"
        window.blazorReceiveRecognitionResult = function(text) {
            DotNet.invokeMethodAsync('你的项目程序集名称', 'ReceiveRecognitionResult', text);
        }
    ");
}

[JSInvokable]
public void ReceiveRecognitionResult(string text)
{
    Console.WriteLine($"识别结果: {text}");
    // 这里可以更新UI或者处理识别文本
}

private async Task StartScreenSharing()
{
    try
    {
        Console.WriteLine("启动屏幕共享...");
        var audioTrackId = await JSRuntime.InvokeAsync<string>("startScreenSharing");
        if (!string.IsNullOrEmpty(audioTrackId))
        {
            isScreenSharing = true;
            StateHasChanged();
            StartTimer();

            // 启动系统音频识别
            await JSRuntime.InvokeVoidAsync("startSystemAudioRecognition", subscriptionKey, region);
        }
        else
        {
            Console.WriteLine("屏幕共享启动失败");
        }
    }
    catch (Exception ex)
    {
        Console.WriteLine($"启动屏幕共享出错: {ex.Message}");
    }
}

// 记得在停止共享时调用停止识别
private async Task StopScreenSharing()
{
    await JSRuntime.InvokeVoidAsync("stopSystemAudioRecognition");
    // 其他停止共享的逻辑...
}
方案B:继续使用.NET Speech SDK(需流转换)

如果你坚持要用.NET版本的SDK,需要把浏览器的MediaStream转换成WAV格式的Stream,再传给Speech SDK:

  1. 添加JS的音频录制函数:
window.startRecordingSystemAudio = async function () {
    const systemAudioStream = window.currentSystemAudioStream;
    if (!systemAudioStream) return false;

    // 创建MediaRecorder录制系统音频为WAV格式
    const mediaRecorder = new MediaRecorder(systemAudioStream, { mimeType: 'audio/wav' });
    const chunks = [];

    mediaRecorder.ondataavailable = (e) => chunks.push(e.data);
    mediaRecorder.onstop = () => {
        // 把录制的Blob转换成ArrayBuffer传给Blazor
        new Blob(chunks, { type: 'audio/wav' }).arrayBuffer().then(buffer => {
            DotNet.invokeMethodAsync('你的项目程序集名称', 'ReceiveAudioBuffer', new Uint8Array(buffer));
        });
    };

    mediaRecorder.start();
    window.currentMediaRecorder = mediaRecorder;
    return true;
};

window.stopRecordingSystemAudio = function () {
    if (window.currentMediaRecorder) {
        window.currentMediaRecorder.stop();
        window.currentMediaRecorder = null;
    }
};
  1. 在Blazor中处理音频流并初始化Speech SDK:
private SpeechRecognizer _recognizer;
private MemoryStream _audioStream;

[JSInvokable]
public async void ReceiveAudioBuffer(byte[] buffer)
{
    if (_audioStream == null)
    {
        _audioStream = new MemoryStream();
    }
    _audioStream.Write(buffer, 0, buffer.Length);
    _audioStream.Position = 0;

    // 初始化Speech SDK
    var config = SpeechConfig.FromSubscription(subscriptionKey, region);
    using var audioConfig = AudioConfig.FromStreamInput(_audioStream);
    _recognizer = new SpeechRecognizer(config, audioConfig);

    _recognizer.Recognizing += (s, e) =>
    {
        if (!string.IsNullOrEmpty(e.Result.Text))
        {
            Console.WriteLine($"识别结果: {e.Result.Text}");
        }
    };

    await _recognizer.StartContinuousRecognitionAsync().ConfigureAwait(false);
}

private async Task StartScreenSharing()
{
    try
    {
        Console.WriteLine("启动屏幕共享...");
        var success = await JSRuntime.InvokeAsync<bool>("startScreenSharing");
        if (success)
        {
            isScreenSharing = true;
            StateHasChanged();
            StartTimer();

            // 启动音频录制
            await JSRuntime.InvokeAsync<bool>("startRecordingSystemAudio");
        }
        else
        {
            Console.WriteLine("屏幕共享启动失败");
        }
    }
    catch (Exception ex)
    {
        Console.WriteLine($"启动屏幕共享出错: {ex.Message}");
    }
}

3. 几个重要的注意事项

  • 一定要删除你之前写的stream.getAudioTracks().forEach(track => track.enabled = false),这会把系统音频也关掉,导致没有输入源。
  • 确保浏览器支持getDisplayMedia的音频捕获(Chrome、Edge默认支持,Firefox需要在about:config中开启media.getDisplayMedia.audio.enabled)。
  • 停止屏幕共享时,务必停止识别和录制,释放所有资源,避免内存泄漏。

备注:内容来源于stack exchange,提问作者Levan Amashukeli

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 18:24:36