You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Librosa实现男女声音互转?男声转女声尝试遇异常求解决

Male-to-Female & Female-to-Male Voice Conversion with Librosa

Hey there! Let's break down why your current code isn't giving you the natural voice conversion you want, and walk through proper, effective ways to do this with Librosa.

Why Your Current Approach Fails

Your code is cropping out a middle chunk of the STFT spectrum (a[10:length-10]), which just strips away a range of frequency components entirely—this is why you're getting weird, distorted audio instead of a natural-sounding female voice.

Male and female voices differ mainly in two core aspects:

  • Fundamental Frequency (F0): Male voices typically sit between 80-160Hz, while female voices range from 160-300Hz.
  • Spectral Envelope: This shapes the "timbre"—female voices tend to have brighter, more prominent high frequencies, while male voices have stronger low-frequency emphasis.

To convert voices naturally, you need to adjust both these elements, not just cut out parts of the spectrum.

Basic Male-to-Female Conversion

Here's a straightforward implementation that adjusts pitch and enhances high frequencies to mimic a female voice:

import librosa
import numpy as np

# Load audio (keep original sample rate with sr=None)
y, sr = librosa.load("/Users/wu4mac/PycharmProjects/SpeechRecognition/weather.wav", sr=None)

# Step 1: Shift pitch up (adjust n_steps based on your audio—4-6 semitones works for most cases)
# Positive n_steps = higher pitch (male → female)
y_pitched = librosa.effects.pitch_shift(y, sr=sr, n_steps=5)

# Step 2: Enhance high frequencies to add brightness
D = librosa.stft(y_pitched)
freqs = librosa.fft_frequencies(sr=sr)
# Boost frequencies above 2000Hz by 20% (tweak this value as needed)
high_freq_indices = np.where(freqs > 2000)[0]
D[high_freq_indices] *= 1.2

# Convert back to audio and save
y_final = librosa.istft(D)
librosa.output.write_wav("male_to_female.wav", y_final, sr)

Basic Female-to-Male Conversion

For the reverse, lower the pitch and boost low frequencies to match male voice characteristics:

# Load your female audio (replace the path with your file)
y_female, sr_female = librosa.load("female_audio.wav", sr=None)

# Step 1: Shift pitch down (negative n_steps)
y_pitched_male = librosa.effects.pitch_shift(y_female, sr=sr_female, n_steps=-5)

# Step 2: Boost low frequencies (below 500Hz)
D_male = librosa.stft(y_pitched_male)
low_freq_indices = np.where(freqs < 500)[0]
D_male[low_freq_indices] *= 1.3

# Save the result
y_final_male = librosa.istft(D_male)
librosa.output.write_wav("female_to_male.wav", y_final_male, sr_female)

Advanced: Pitch-Shifting with Precise F0 Control

For more natural results, you can extract the exact fundamental frequency (F0) of the original voice and scale it to match the target gender. This uses Librosa's pyin for accurate F0 detection and the phase vocoder to shift frequencies smoothly without distorting timing:

import librosa
import numpy as np

y, sr = librosa.load("/Users/wu4mac/PycharmProjects/SpeechRecognition/weather.wav", sr=None)

# Extract F0 (fundamental frequency)
f0, voiced_flag, _ = librosa.pyin(y, fmin=librosa.note_to_hz('C2'), fmax=librosa.note_to_hz('C7'))

# Scale F0 for male → female (multiply by ~1.4-1.6; adjust based on your audio)
target_f0 = f0 * 1.5
# Keep unvoiced segments unchanged
target_f0[~voiced_flag] = f0[~voiced_flag]

# Calculate frequency shift factor
shift_factor = target_f0 / f0
# Handle NaN values (unvoiced parts)
shift_factor[np.isnan(shift_factor)] = 1.0

# Apply phase vocoder to shift frequencies smoothly
D = librosa.stft(y)
D_shifted = librosa.phase_vocoder(D, shift_factor, hop_length=512)

# Convert back to audio
y_final_advanced = librosa.istft(D_shifted)
librosa.output.write_wav("male_to_female_advanced.wav", y_final_advanced, sr)

Key Tips

  • Adjust n_steps or the F0 scaling factor based on your specific audio—some voices may need more/less pitch adjustment to sound natural.
  • Tweak the frequency boost values (1.2, 1.3) to fine-tune the timbre to your liking.
  • For even more polished results, you could explore advanced techniques like voice conversion with GANs, but that's beyond basic Librosa usage.

内容的提问来源于stack exchange,提问作者Shawn Plus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 08:22:04