You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Essentia.js检测测试音频特征结果不一致问题排查

问题描述

我需要检测音频文件中的音符位置、音高等特征,选用Essentia.js库实现,具体代码如下:

const Essentia = require('essentia.js');
const fs = require('fs');
const glob = require('glob');
const path = require('path');
const wav = require('node-wav');

const essentia = new Essentia.Essentia(Essentia.EssentiaWASM);
const audioDir = path.join('test', 'audio', '**', '*.wav');
const audioPaths = glob.globSync(audioDir);
const results = [];

// 遍历文件夹中的每个文件并检测音频特征
audioPaths.forEach((audioPath) => {
  console.log(`Analyzing ${audioPath}`);
  const fileBuffer = fs.readFileSync(audioPath);
  const audioBuffer = wav.decode(fileBuffer);
  const audioVector = essentia.arrayToVector(audioBuffer.channelData[0]);
  const melodia = essentia.PredominantPitchMelodia(audioVector).pitch;
  const segments = essentia.PitchContourSegmentation(melodia, audioVector);
  results.push({
    audioPath,
    durations: essentia.vectorToArray(segments.duration),
    onsets: essentia.vectorToArray(segments.onset),
    pitches: essentia.vectorToArray(segments.MIDIpitch)
  });
});

// 并排输出属性用于对比
results.forEach(result => console.log('durations', result.audioPath, result.durations));
results.forEach(result => console.log('onsets', result.audioPath, result.onsets));
results.forEach(result => console.log('pitches', result.audioPath, result.pitches));

测试音频文件包含相同数量的音符,但检测结果不一致,不同文件的音符检测数量、时长、起始点(Onsets)、音高结果均有差异:

时长结果

durations test/audio/velocity-sin.wav Float32Array(10) [
  0.5224489569664001,
  0.5369614362716675,
  0.5369614362716675,
  0.5369614362716675,
  0.5369614362716675,
  0.10158730298280716,
  0.1160997748374939,
  0.14222222566604614,
  0.09868481010198593,
  0.10739228874444962
]
durations test/audio/velocity-saw.wav Float32Array(7) [
  0.5195465087890625,
  0.528253972530365,
  0.528253972530365,
  0.528253972530365,
  0.528253972530365,
  0.12480725347995758,
  0.15673469007015228
]
durations test/audio/velocity-piano.wav Float32Array(3) [
  2.983764171600342,
  2.002721071243286,
  3.0040817260742188
]

起始点(Onsets)结果

onsets test/audio/scale-square.wav Float32Array(12) [
  0,
  0.9839455485343933,
  1.9824036359786987,
  2.9808616638183594,
  3.9851248264312744,
  4.983582973480225,
  5.604716777801514,
  5.979138374328613,
  6.65541934967041,
  6.759909152984619,
  6.977596282958984,
  7.7380499839782715
]
onsets test/audio/scale-sin.wav Float32Array(11) [
  0,
  0.9839455485343933,
  1.9824036359786987,
  2.977959156036377,
  3.982222318649292,
  4.980680465698242,
  5.784671306610107,
  5.979138374328613,
  6.977596282958984,
  7.5290703773498535,
  7.679999828338623
]
onsets test/audio/scale-saw.wav Float32Array(10) [
  0,
  0.9868480563163757,
  1.9853061437606812,
  2.983764171600342,
  3.558458089828491,
  3.689070224761963,
  3.9851248264312744,
  4.97777795791626,
  5.982040882110596,
  6.980498790740967
]

音高结果

pitches test/audio/scale2-sin.wav Float32Array(21) [
  60, 61, 62, 63, 64, 65, 66,
  67, 68, 69, 70, 71, 86, 72,
  92, 90, 91, 88, 93, 93, 90
]
pitches test/audio/scale2-saw.wav Float32Array(19) [
  60, 61, 62, 63, 64, 65, 66,
  67, 68, 69, 70, 71, 72, 92,
  91, 89, 92, 88, 92
]
pitches test/audio/scale2-piano.wav Float32Array(11) [
  60, 61, 62, 63, 64,
  65, 66, 67, 68, 69,
  70
]

请问该现象的原因是什么?是Essentia.js的bug还是我的实现存在问题?


解答

这种差异既不是Essentia.js的bug,也不是你的代码实现问题,核心原因来自音频本身的特性和算法的固有局限性:

  1. 音频波形的谐波结构差异
    你测试的正弦波、锯齿波、钢琴音频,谐波含量差异极大:

    • 正弦波只有基频,无谐波,音高检测精准,但音符衰减平滑,容易被算法拆成多个短片段;
    • 锯齿波含丰富奇次谐波,能强化音高特征,但谐波会干扰起始点检测,导致部分边界被合并;
    • 钢琴属于冲击力强的原声乐器,音符起始有清晰瞬态,衰减阶段泛音变化自然,算法更容易识别连续长音符,尾音会被合并到主音符中。
  2. PitchContourSegmentation的参数敏感性
    你调用PitchContourSegmentation时用了默认参数,分割阈值、平滑窗口等参数直接影响结果:

    • 对平滑衰减的正弦波,默认阈值会把缓慢变化的音高轮廓拆成多个小片段;
    • 对谐波丰富的锯齿波,算法会把相近音高轮廓合并,减少检测片段数;
    • 钢琴音瞬态清晰,算法能识别明确边界,所以检测片段数最少。
  3. PredominantPitchMelodia的音高检测误差
    该算法基于旋律检测音高,对单旋律表现较好,但不同波形的音高稳定性不同:

    • 正弦波音高绝对稳定,但静音/衰减阶段易出现误检测(比如你看到的86、92这类异常MIDI音高);
    • 锯齿波的谐波会导致音高微小偏移,但整体更接近真实值;
    • 钢琴泛音结构复杂,但算法能更好追踪主音高,异常值最少。
  4. 音符衰减特性影响时长计算
    不同乐器的音符衰减速度不同:

    • 正弦波衰减慢,算法会把衰减过程中的小变化当成独立片段,导致时长碎片化;
    • 钢琴衰减快,尾音迅速消失,算法直接将整个音符视为长片段,时长更接近真实值。

优化建议

若要获得更一致的检测结果,可做以下调整:

  • 手动配置PredominantPitchMelodia的minFrequency和maxFrequency参数,限定检测范围,减少异常音高;
  • 调整PitchContourSegmentation的segmentationThreshold参数,针对不同波形设置合适阈值,比如给正弦波提高阈值,减少不必要分割;
  • 对音频做预处理,添加硬阈值门限,过滤衰减到一定程度的静音部分,避免算法误判。

内容的提问来源于stack exchange,提问作者Kim T

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 04:02:03