如何用Librosa实现男女声音互转?男声转女声尝试遇异常求解决
Hey there! Let's break down why your current code isn't giving you the natural voice conversion you want, and walk through proper, effective ways to do this with Librosa.
Why Your Current Approach Fails
Your code is cropping out a middle chunk of the STFT spectrum (a[10:length-10]), which just strips away a range of frequency components entirely—this is why you're getting weird, distorted audio instead of a natural-sounding female voice.
Male and female voices differ mainly in two core aspects:
- Fundamental Frequency (F0): Male voices typically sit between 80-160Hz, while female voices range from 160-300Hz.
- Spectral Envelope: This shapes the "timbre"—female voices tend to have brighter, more prominent high frequencies, while male voices have stronger low-frequency emphasis.
To convert voices naturally, you need to adjust both these elements, not just cut out parts of the spectrum.
Basic Male-to-Female Conversion
Here's a straightforward implementation that adjusts pitch and enhances high frequencies to mimic a female voice:
import librosa import numpy as np # Load audio (keep original sample rate with sr=None) y, sr = librosa.load("/Users/wu4mac/PycharmProjects/SpeechRecognition/weather.wav", sr=None) # Step 1: Shift pitch up (adjust n_steps based on your audio—4-6 semitones works for most cases) # Positive n_steps = higher pitch (male → female) y_pitched = librosa.effects.pitch_shift(y, sr=sr, n_steps=5) # Step 2: Enhance high frequencies to add brightness D = librosa.stft(y_pitched) freqs = librosa.fft_frequencies(sr=sr) # Boost frequencies above 2000Hz by 20% (tweak this value as needed) high_freq_indices = np.where(freqs > 2000)[0] D[high_freq_indices] *= 1.2 # Convert back to audio and save y_final = librosa.istft(D) librosa.output.write_wav("male_to_female.wav", y_final, sr)
Basic Female-to-Male Conversion
For the reverse, lower the pitch and boost low frequencies to match male voice characteristics:
# Load your female audio (replace the path with your file) y_female, sr_female = librosa.load("female_audio.wav", sr=None) # Step 1: Shift pitch down (negative n_steps) y_pitched_male = librosa.effects.pitch_shift(y_female, sr=sr_female, n_steps=-5) # Step 2: Boost low frequencies (below 500Hz) D_male = librosa.stft(y_pitched_male) low_freq_indices = np.where(freqs < 500)[0] D_male[low_freq_indices] *= 1.3 # Save the result y_final_male = librosa.istft(D_male) librosa.output.write_wav("female_to_male.wav", y_final_male, sr_female)
Advanced: Pitch-Shifting with Precise F0 Control
For more natural results, you can extract the exact fundamental frequency (F0) of the original voice and scale it to match the target gender. This uses Librosa's pyin for accurate F0 detection and the phase vocoder to shift frequencies smoothly without distorting timing:
import librosa import numpy as np y, sr = librosa.load("/Users/wu4mac/PycharmProjects/SpeechRecognition/weather.wav", sr=None) # Extract F0 (fundamental frequency) f0, voiced_flag, _ = librosa.pyin(y, fmin=librosa.note_to_hz('C2'), fmax=librosa.note_to_hz('C7')) # Scale F0 for male → female (multiply by ~1.4-1.6; adjust based on your audio) target_f0 = f0 * 1.5 # Keep unvoiced segments unchanged target_f0[~voiced_flag] = f0[~voiced_flag] # Calculate frequency shift factor shift_factor = target_f0 / f0 # Handle NaN values (unvoiced parts) shift_factor[np.isnan(shift_factor)] = 1.0 # Apply phase vocoder to shift frequencies smoothly D = librosa.stft(y) D_shifted = librosa.phase_vocoder(D, shift_factor, hop_length=512) # Convert back to audio y_final_advanced = librosa.istft(D_shifted) librosa.output.write_wav("male_to_female_advanced.wav", y_final_advanced, sr)
Key Tips
- Adjust
n_stepsor the F0 scaling factor based on your specific audio—some voices may need more/less pitch adjustment to sound natural. - Tweak the frequency boost values (1.2, 1.3) to fine-tune the timbre to your liking.
- For even more polished results, you could explore advanced techniques like voice conversion with GANs, but that's beyond basic Librosa usage.
内容的提问来源于stack exchange,提问作者Shawn Plus

