Python环境下WAV文件分帧FFT计算及补零、归一化技术咨询
Hey there! I'll walk you through each of your questions with practical, beginner-friendly code examples—since you're new to Python and working on GCP, I’ll keep this straightforward and tailored to your workflow.
1. Splitting the WAV File into Frames Matching 30fps
First, let’s calculate the frame size: at 48kHz sample rate and 30fps, each frame needs to cover 48000 / 30 = 1600 audio samples.
You can split the audio array into consecutive frames directly with NumPy slicing. If you’re using librosa, it has built-in frame-handling tools, but manual slicing is easier to follow when you’re starting out.
Here’s how to do it:
import numpy as np from scipy.io import wavfile # Load your WAV file (see GCP note below for cloud storage access) sample_rate, audio_data = wavfile.read("your_audio.wav") # Convert stereo to mono if needed (most audio processing uses mono for this task) if len(audio_data.shape) > 1: audio_data = np.mean(audio_data, axis=1) frame_size = int(sample_rate / 30) # 1600 samples per frame total_frames = len(audio_data) // frame_size # Split into frames (we'll discard leftover samples if the audio length isn't a perfect multiple of frame size) frames = np.array([audio_data[i*frame_size : (i+1)*frame_size] for i in range(total_frames)])
GCP Tip: If your WAV is stored in Google Cloud Storage (GCS), use the google-cloud-storage library to download it to a temporary file first:
from google.cloud import storage import tempfile storage_client = storage.Client() bucket = storage_client.bucket("your-bucket-name") blob = bucket.blob("path/to/your/audio.wav") # Download to a temporary file (avoids saving to permanent storage) with tempfile.NamedTemporaryFile(suffix=".wav", delete=False) as tmp: blob.download_to_file(tmp) tmp_path = tmp.name # Now load the audio from the temporary path sample_rate, audio_data = wavfile.read(tmp_path)
2. Padding Frames to the Next Power of Two (2048 samples)
To pad each 1600-sample frame to 2048 (the next power of two after 1600), you have two easy options:
Option 1: Manual padding with NumPy
target_length = 2048 # 2^11, the smallest power of two larger than 1600 padded_frames = np.array([np.pad(frame, (0, target_length - len(frame)), mode='constant') for frame in frames])
Option 2: Automatic padding in FFT (more efficient)
When you call scipy.fft.fft(), you can specify the n parameter to tell it to pad the frame to your target length automatically. This saves you from creating padded frames explicitly:
from scipy.fft import fft # This will auto-pad each frame to 2048 samples during FFT computation fft_results = np.array([fft(frame, n=2048) for frame in frames])
3. Normalizing FFT Values to the [-1, 1] Range
FFT outputs complex numbers, so we’ll handle the most common use case first: normalizing the magnitude of the FFT values. If you need to work with raw complex components (real/imaginary parts), we’ll cover that too.
For FFT Magnitudes (most common for audio analysis)
# Get the magnitude of each FFT result fft_magnitudes = np.abs(fft_results) # Normalize to [-1, 1]: first scale to [0,1], then shift to [-1,1] max_magnitude = np.max(fft_magnitudes) normalized_fft = (fft_magnitudes / max_magnitude) * 2 - 1
For Raw Complex FFT Values
If you need to preserve the real and imaginary components, normalize each part separately:
# Split real and imaginary parts real_parts = fft_results.real imag_parts = fft_results.imag # Normalize real parts to [-1,1] max_real = np.max(np.abs(real_parts)) normalized_real = real_parts / max_real # Normalize imaginary parts to [-1,1] max_imag = np.max(np.abs(imag_parts)) normalized_imag = imag_parts / max_imag # Combine back into a complex array (or write real/imag values separately to text) normalized_fft_complex = normalized_real + 1j * normalized_imag
Writing Results to a Text File
Finally, write each frame’s normalized FFT values to a text file, with one frame per line:
with open("fft_output.txt", "w") as f: for frame_fft in normalized_fft: # Convert the array to a comma-separated string line = ",".join(map(str, frame_fft)) + "\n" f.write(line)
GCP Tip: Upload the output file back to GCS when you’re done:
output_blob = bucket.blob("path/to/output/fft_output.txt") output_blob.upload_from_filename("fft_output.txt")
内容的提问来源于stack exchange,提问作者Yadhu

