You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将AnalyserNode连接到SpeechSynthesisUtterance实现语音可视化?

实现SpeechSynthesisUtterance的语音可视化

可以实现语音可视化,但无法直接将AnalyserNode连接到SpeechSynthesisUtterance的输出——因为这个Web API没有暴露直接访问其音频流的节点接口。不过有两种可行的方案来实现需求:

方案1:捕获系统音频流(跨浏览器可靠方案)

这种方法通过捕获系统的整体音频输出来获取SpeechSynthesis的语音,需要用户授权,步骤如下:

  1. 请求系统音频流权限
async function getSystemAudioStream() {
  try {
    const stream = await navigator.mediaDevices.getDisplayMedia({
      video: false,
      audio: true
    });
    return stream;
  } catch (err) {
    console.error('获取系统音频流失败:', err);
    throw err;
  }
}
  1. 接入Web Audio API并连接AnalyserNode
// 初始化AudioContext和AnalyserNode
const audioCtx = new (window.AudioContext || window.webkitAudioContext)();
const analyser = audioCtx.createAnalyser();
analyser.fftSize = 2048;
const bufferLength = analyser.frequencyBinCount;
const dataArray = new Uint8Array(bufferLength);

// 获取系统音频流并创建媒体源节点
const systemStream = await getSystemAudioStream();
const sourceNode = audioCtx.createMediaStreamSource(systemStream);

// 连接节点链路:系统音频源 → Analyser → 扬声器
sourceNode.connect(analyser);
analyser.connect(audioCtx.destination);

// 实时绘制可视化效果(示例用Canvas)
function renderVisualization(canvas) {
  const ctx = canvas.getContext('2d');
  canvas.width = window.innerWidth;
  canvas.height = window.innerHeight;

  function draw() {
    requestAnimationFrame(draw);
    analyser.getByteFrequencyData(dataArray);

    ctx.fillStyle = '#000';
    ctx.fillRect(0, 0, canvas.width, canvas.height);

    const barWidth = (canvas.width / bufferLength) * 2.5;
    let x = 0;
    for (let i = 0; i < bufferLength; i++) {
      const barHeight = (dataArray[i] / 255) * canvas.height;
      ctx.fillStyle = `rgb(${barHeight + 100}, 50, 50)`;
      ctx.fillRect(x, canvas.height - barHeight, barWidth, barHeight);
      x += barWidth + 1;
    }
  }
  draw();
}

// 调用可视化函数(传入你的Canvas元素)
renderVisualization(document.getElementById('visualizer'));
  1. 播放文本转语音
const utterance = new SpeechSynthesisUtterance('这是一段测试语音');
window.speechSynthesis.speak(utterance);

注意:该方案会捕获所有系统音频输出,不仅仅是SpeechSynthesis的语音,需要用户在授权弹窗中选择“系统音频”。

方案2:MediaStreamDestination中转(兼容性有限)

部分Chromium系浏览器支持非标准的utterance.audioDestination属性,可以将SpeechSynthesis的输出直接路由到MediaStreamDestination,再接入AnalyserNode:

const audioCtx = new (window.AudioContext || window.webkitAudioContext)();
const analyser = audioCtx.createAnalyser();
analyser.fftSize = 2048;
const bufferLength = analyser.frequencyBinCount;
const dataArray = new Uint8Array(bufferLength);

// 创建媒体流目标节点
const mediaStreamDest = audioCtx.createMediaStreamDestination();
analyser.connect(audioCtx.destination);

// 创建并配置语音合成实例
const utterance = new SpeechSynthesisUtterance('测试语音可视化');
// 仅部分浏览器支持该属性
utterance.audioDestination = mediaStreamDest.stream;

// 将媒体流目标接入AnalyserNode
const sourceNode = audioCtx.createMediaStreamSource(mediaStreamDest.stream);
sourceNode.connect(analyser);

// 播放语音
window.speechSynthesis.speak(utterance);

// 可视化绘制逻辑同方案1
function renderVisualization(canvas) {
  // 同方案1的绘制代码
}

注意:audioDestination并非标准API,兼容性较差,不推荐在生产环境使用。

总结

如果需要跨浏览器的可靠方案,优先选择方案1;若仅针对特定浏览器环境,可以尝试方案2。另外,也可以考虑使用第三方TTS服务(直接返回音频流),这种方式更灵活,可直接将流接入AnalyserNode。

内容的提问来源于stack exchange,提问作者mustafa.salaheldin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 20:17:09