You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java提取WAV文件基频求助:人声元音音素检测方案及代码示例

Hey there! I get that diving into audio processing can feel overwhelming at first, but extracting the fundamental frequency (F0) of vowel phonemes is a totally achievable task—especially since vowels have strong periodicity, which makes them perfect for refined autocorrelation-based methods like the YIN algorithm. Let's walk through the key ideas and a practical Java implementation.

Core Approach: Why YIN Over Basic Autocorrelation?

You mentioned trying autocorrelation and FFT—basic autocorrelation can work, but it often picks up harmonic peaks instead of the true fundamental. The YIN algorithm fixes this by calculating a "difference function" that emphasizes the true periodicity of the signal, making it far more robust for speech (especially vowels). Here's the high-level workflow:

  • Preprocess the audio: Split into short frames (since speech changes over time), apply a window function to reduce spectral leakage, and remove any DC offset.
  • Compute the difference function: Measures how similar the signal is to itself when shifted by different delays.
  • Normalize and threshold: Refine the difference function to eliminate false peaks.
  • Find the fundamental period: Locate the first valid peak in the refined function, then convert it to frequency using the sample rate.
  • Post-process: Smooth the F0 values to remove outliers (vowels are stable, so sudden jumps are likely errors).
Java Implementation (Simplified YIN Algorithm)

This snippet assumes you have a raw audio buffer (float array) of your vowel sample, sampled at 44100 Hz. We'll process it frame by frame to get F0 values over time.

import java.util.Arrays;

public class VowelF0Extractor {
    private static final int SAMPLE_RATE = 44100;
    private static final int FRAME_SIZE = 2048; // Adjust based on your needs
    private static final int HOP_SIZE = 1024; // 50% overlap between frames
    private static final double YIN_THRESHOLD = 0.15; // Sensitivity for peak detection

    public static void main(String[] args) {
        // Replace this with your actual vowel audio data (float array, normalized to [-1, 1])
        float[] audioBuffer = loadVowelAudio();
        
        // Process each frame to get F0 values
        for (int i = 0; i + FRAME_SIZE <= audioBuffer.length; i += HOP_SIZE) {
            float[] frame = Arrays.copyOfRange(audioBuffer, i, i + FRAME_SIZE);
            double f0 = calculateF0(frame);
            System.out.printf("Frame %d: F0 = %.2f Hz%n", i / HOP_SIZE, f0);
        }
    }

    private static double calculateF0(float[] frame) {
        // Step 1: Preprocess frame - remove DC offset and apply Hamming window
        removeDCOffset(frame);
        applyHammingWindow(frame);

        // Step 2: Compute YIN difference function
        int maxDelay = frame.length / 2;
        double[] difference = new double[maxDelay];
        for (int tau = 1; tau < maxDelay; tau++) {
            double sum = 0;
            for (int i = 0; i < maxDelay; i++) {
                double diff = frame[i] - frame[i + tau];
                sum += diff * diff;
            }
            difference[tau] = sum;
        }

        // Step 3: Compute cumulative mean normalized difference function (CMNDF)
        double[] cmndf = new double[maxDelay];
        cmndf[0] = 1;
        double cumulativeSum = 0;
        for (int tau = 1; tau < maxDelay; tau++) {
            cumulativeSum += difference[tau];
            cmndf[tau] = difference[tau] / ((cumulativeSum / tau) * tau);
        }

        // Step 4: Find the first tau where CMNDF crosses the threshold
        int fundamentalTau = -1;
        for (int tau = 1; tau < maxDelay - 1; tau++) {
            if (cmndf[tau] < YIN_THRESHOLD && cmndf[tau] < cmndf[tau + 1]) {
                // Interpolate to get a more precise tau (optional but improves accuracy)
                fundamentalTau = tau + (int) ((cmndf[tau] - cmndf[tau + 1]) / (cmndf[tau - 1] - 2 * cmndf[tau] + cmndf[tau + 1]));
                break;
            }
        }

        // If no valid tau found, return 0 (or handle silence/non-periodic signal)
        if (fundamentalTau == -1) {
            return 0;
        }

        // Convert tau to frequency: F0 = sample rate / tau
        return SAMPLE_RATE / (double) fundamentalTau;
    }

    private static void removeDCOffset(float[] frame) {
        double mean = 0;
        for (float sample : frame) {
            mean += sample;
        }
        mean /= frame.length;
        for (int i = 0; i < frame.length; i++) {
            frame[i] -= mean;
        }
    }

    private static void applyHammingWindow(float[] frame) {
        for (int i = 0; i < frame.length; i++) {
            frame[i] *= 0.54 - 0.46 * Math.cos(2 * Math.PI * i / (frame.length - 1));
        }
    }

    // Dummy method to load audio - replace with your actual audio loading logic
    private static float[] loadVowelAudio() {
        // For testing, generate a synthetic vowel-like sine wave (100 Hz)
        float[] synthetic = new float[FRAME_SIZE * 4];
        for (int i = 0; i < synthetic.length; i++) {
            synthetic[i] = (float) Math.sin(2 * Math.PI * 100 * i / SAMPLE_RATE);
        }
        return synthetic;
    }
}
Key Notes for Your Project
  • Audio Loading: Replace the loadVowelAudio() method with code to read your actual audio file (WAV is easiest—use Java's AudioInputStream to extract raw samples).
  • Parameter Tuning: Adjust FRAME_SIZE, HOP_SIZE, and YIN_THRESHOLD based on your audio:
    • Smaller frames (e.g., 1024) give better time resolution but less frequency accuracy.
    • Lower threshold values make the detector more sensitive but may pick up noise.
  • Post-Processing: Add a moving average or median filter to smooth F0 values—vowels should have a stable F0, so sudden spikes are almost always errors.

内容的提问来源于stack exchange,提问作者Sayan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:52:04