Python3中scikits.talkbox的lpc无法使用,求替代方案及共振峰频率估算示例
Ah, I feel your pain—scikits.talkbox is basically abandoned for Python 3, so that LPC function is definitely not reliable these days. Let’s walk through some solid replacements and a complete example to estimate formants in Python 3 that actually works.
scikits.talkbox.lpc 1. Use Librosa's LPC Implementation
Librosa is a well-maintained, widely used audio processing library that fully supports Python 3. Its librosa.lpc function is a drop-in replacement for the old scikits version.
First, install Librosa if you haven't:
pip install librosa
Then use it like this:
import librosa # y = your audio signal array, n_order = desired LPC order lpc_coeffs = librosa.lpc(y, order=n_order)
2. Use python_speech_features (Lightweight Option)
If you prefer a library focused specifically on speech features, python_speech_features also has a reliable LPC implementation:
Install it first:
pip install python_speech_features
Call the function like so:
from python_speech_features import lpc # signal = your audio signal, numcep = LPC order lpc_coeffs = lpc(signal, numcep=n_order)
Formants are the resonant frequencies of the vocal tract, and we can calculate them from LPC coefficients by finding roots of the LPC polynomial. Here's a full, working example using Librosa:
import librosa import numpy as np def estimate_formants(y, sr, lpc_order=12): # Optional: Apply pre-emphasis to boost high frequencies (reduces noise impact) y_preemphasized = librosa.effects.preemphasis(y) # Extract LPC coefficients lpc_coeffs = librosa.lpc(y_preemphasized, order=lpc_order) # Find roots of the LPC polynomial roots = np.roots(lpc_coeffs) # Keep only roots in the upper half of the complex plane (correspond to positive frequencies) roots = [root for root in roots if np.imag(root) > 0] # Convert roots to frequencies (Hz) angles = np.arctan2(np.imag(roots), np.real(roots)) frequencies = angles * (sr / (2 * np.pi)) # Calculate magnitude of each root (indicates formant strength) magnitudes = 1 / np.abs(roots) # Filter out low-frequency noise (ignore values below 50Hz) valid_mask = frequencies > 50 frequencies = frequencies[valid_mask] magnitudes = magnitudes[valid_mask] # Sort formants by frequency (and match magnitudes to the sorted order) sorted_indices = np.argsort(frequencies) frequencies = frequencies[sorted_indices] magnitudes = magnitudes[sorted_indices] # Return top 4 formants (F1-F4, the most useful for speech analysis) return frequencies[:4], magnitudes[:4] # Example usage # Load your audio file (replace with your file path) y, sr = librosa.load("your_audio.wav", sr=None) # For better accuracy, process audio in frames (this example uses the whole signal for simplicity) formant_freqs, formant_mags = estimate_formants(y, sr) print(f"Estimated formants (Hz): {formant_freqs.round(2)}") print(f"Formant magnitudes: {formant_mags.round(2)}")
Quick Tips for Better Results
- LPC Order: A good rule of thumb is
2 + (sr // 1000)—for 16kHz audio, use 18; for 8kHz, use 10. Adjust based on your specific audio. - Frame Processing: In real-world use, split the audio into 20-30ms frames with overlap, then estimate formants per frame (Librosa's
librosa.util.framecan help with this). - Noise Reduction: If your audio is noisy, add a low-pass filter or use noise reduction techniques before estimating formants.
内容的提问来源于stack exchange,提问作者david

