You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Web Audio API的MIDI播放器播放时出现点击声问题排查

问题

我尝试用Web Audio API构建MIDI播放器,通过Tone.js将MIDI文件解析为JSON格式,使用MP3文件作为音符采样播放。代码能正常播放MIDI文件,但播放过程中存在点击声,且速度(tempo)越高点击声越明显,即使低至50bpm也无法完全消除。

测试用MIDI文件为C大调音阶,经timidity验证无异常。相关代码片段如下:

采样加载与播放

// 创建音频采样
static async setupSample(audioContext, filepath) {
    const response = await fetch(filepath);
    const arrayBuffer = await response.arrayBuffer();
    const audioBuffer = await audioContext.decodeAudioData(arrayBuffer);
    return audioBuffer;
}     

// 播放单个采样
static playSample(audioContext, audioBuffer, time) {
    const sampleSource = new AudioBufferSourceNode(audioContext, {
        buffer: audioBuffer,
        playbackRate: 1,
    });
    sampleSource.connect(audioContext.destination);
    sampleSource.start(time);
    return sampleSource;
}   

采样调度逻辑

async start() {
    this.startTime = this.audioCtx.currentTime;
    this.play();
}   

play() {
    let nextNote = this.notes[this.noteIndex];

    // 调度采样
    while ((nextNote.time + this.startTime) - this.audioCtx.currentTime <= 0.250) {
        let s = Audio.playSample(this.audioCtx, this.samples[nextNote.midi], this.startTime + nextNote.time);
        s.stop(this.startTime + nextNote.time + nextNote.duration);

        this.noteIndex++;
        if (this.noteIndex == this.notes.length) {
            break;
        }

        nextNote = this.notes[this.noteIndex];
    }

    if (this.noteIndex == this.notes.length) {
        return;
    }

    requestAnimationFrame(() => {
        this.play();
    });
}       
原因分析与解决方案

核心原因

  1. 硬截断音频导致波形突变:直接调用stop()会瞬间切断音频流,若截断点不在音频波形的零交叉点(波形正负转换的过零点),就会产生尖锐的点击声。MP3压缩可能在采样开头/结尾引入微小瞬态,进一步加剧问题。
  2. 采样播放速率未匹配MIDI音高:当前playbackRate固定为1,但MIDI音符的音高和采样原始音高大概率不匹配,导致采样时长与音符实际时长偏差,截断时更易出现突兀的波形变化。
  3. 调度精度不足:requestAnimationFrame基于视觉帧率(约60fps,间隔16ms)调度,而音频时钟需要毫秒级甚至微秒级精度。当tempo升高、音符间隔变小时,这种精度偏差会导致start/stop时间与实际音频播放不同步,引发点击。

解决方案

1. 添加淡入淡出避免硬截断

修改playSample函数,通过GainNode实现快速淡入淡出,确保音频平滑启动和结束:

static playSample(audioContext, audioBuffer, time, duration) {
    const sampleSource = new AudioBufferSourceNode(audioContext, {
        buffer: audioBuffer,
    });
    const gainNode = new GainNode(audioContext, { gain: 0 });
    
    sampleSource.connect(gainNode);
    gainNode.connect(audioContext.destination);

    // 5ms淡入,消除开头瞬态
    gainNode.gain.setValueAtTime(0, time);
    gainNode.gain.linearRampToValueAtTime(1, time + 0.005);

    // 10ms淡出,在音符结束前降低音量,避免硬截断
    const stopTime = time + duration;
    gainNode.gain.linearRampToValueAtTime(0, stopTime - 0.01);

    sampleSource.start(time);
    sampleSource.stop(stopTime);
    return sampleSource;
}

调用时传入音符时长:

let s = Audio.playSample(
    this.audioCtx, 
    this.samples[nextNote.midi], 
    this.startTime + nextNote.time,
    nextNote.duration
);

2. 匹配采样播放速率与MIDI音高

假设采样原始音高为C4(MIDI编号60),根据当前MIDI音符计算正确的playbackRate:

// 采样原始MIDI音高基准
const baseMidi = 60;
const playbackRate = Math.pow(2, (nextNote.midi - baseMidi) / 12);

// 创建AudioBufferSourceNode时设置速率
const sampleSource = new AudioBufferSourceNode(audioContext, {
    buffer: audioBuffer,
    playbackRate: playbackRate,
});

这不仅能修正音高,还能让采样时长与音符时长更匹配,减少截断突兀感。

3. 使用音频专用调度方式

替换requestAnimationFrame为基于音频时钟的调度,比如用setTimeout结合音频上下文currentTime计算延迟:

play() {
    let nextNote = this.notes[this.noteIndex];
    const scheduleTime = nextNote.time + this.startTime;
    const delay = scheduleTime - this.audioCtx.currentTime;

    if (delay <= 0.250) {
        // 调度当前音符
        let s = Audio.playSample(
            this.audioCtx, 
            this.samples[nextNote.midi], 
            scheduleTime,
            nextNote.duration
        );

        this.noteIndex++;
        if (this.noteIndex < this.notes.length) {
            // 提前调度下一个音符
            const nextScheduleTime = this.notes[this.noteIndex].time + this.startTime;
            setTimeout(() => this.play(), (nextScheduleTime - this.audioCtx.currentTime) * 1000);
        }
    } else {
        // 延迟到接近调度时间再执行
        setTimeout(() => this.play(), delay * 1000);
    }
}

4. 优化采样文件

  • 将MP3采样转换为WAV格式:无压缩的WAV波形更稳定,避免MP3解码带来的微小波形突变。
  • 预处理采样:用音频编辑工具(如Audacity)修剪采样开头/结尾的多余瞬态,确保开头和结尾在零交叉点附近。

内容的提问来源于stack exchange,提问作者kargirwar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 09:25:44