关于离散余弦变换(DCT)输出理解及音频压缩的技术咨询
Hey there, let's unpack why your current setup isn't working as expected—this is a super common gotcha when diving into frequency-domain audio processing, so you're not alone here!
First, let's break down the core misunderstandings that might be throwing things off:
1. You're using global DCT on too-long segments
A 5-second audio clip is way too long to run a single DCT on. Audio is a time-varying signal—the frequency content changes constantly (think a vocal line shifting to a drum hit, or a guitar chord changing). A global DCT over 5 seconds will average out all those time-dependent frequency features, so you'll end up with a "blurry" frequency profile that doesn't represent what's actually happening at any given moment in the audio. This makes it impossible to pick truly "important" frequencies that matter for the actual sound.
2. "Important frequencies" aren't one-size-fits-all
Trying to use a single set of "most important frequencies" across all your 5-second segments is a mistake. Different audio segments (e.g., a bass-heavy drum loop vs. a high-pitched vocal) have entirely different critical frequency bands. A frequency that's irrelevant in one segment might be the most important for the clarity of another.
3. DCT alone isn't optimized for audio compression
While DCT is great for static signals like images, raw DCT lacks features needed for audio:
- No handling of frame boundary artifacts: When you slice audio into segments and apply DCT, the edges between segments can cause audible pops or distortion (similar to the Gibbs phenomenon in Fourier transforms).
- Doesn't account for human auditory perception: Your current approach is probably picking frequencies with the highest energy, but human ears don't perceive all frequencies equally. We're far more sensitive to 2-4kHz (speech, midrange instruments) than we are to very low or very high frequencies—so wasting bits on high-energy but inaudible frequencies is inefficient, and missing low-energy but perceptually critical ones ruins sound quality.
Fixes to Get Your Compression Working
Here's how to adjust your approach to align with how audio compression actually works:
Switch to short-time frame processing
Instead of 5-second segments, slice your audio into short, overlapping frames (20-50ms is standard, e.g., 20ms at 44.1kHz is 882 samples). For each frame:
- Apply a window function (like a Hamming window) to reduce boundary artifacts.
- Use either Short-Time DCT (SDCT) or, more commonly in audio compression, Modified DCT (MDCT)—MDCT is used in standards like MP3 and AAC because it uses overlapping frames to eliminate boundary distortion entirely.
Define "important" based on human hearing, not just energy
Use a perceptual weighting filter (like the A-weighting curve, or more advanced models from psychoacoustics) to prioritize frequencies that humans actually care about. For each frame, rank frequency components by their perceptual importance, not just their raw energy. This lets you discard more components without hurting audio quality.
Process each frame independently
Don't try to use a global set of "top frequencies" across all segments. Each short frame has its own unique frequency profile—keep the top N perceptually important components per frame, not across the entire audio file.
Example revised workflow
Original audio → Slice into 20ms overlapping frames (Hamming window applied) → Run MDCT on each frame → Apply perceptual weighting to rank frequency components → Keep top 20-30% of components per frame → Store frame metadata + retained components → For decoding: Reconstruct each frame via inverse MDCT → Overlap-add frames to eliminate gaps
This approach aligns with how modern audio compression works, and should fix the issues you're seeing with your initial DCT setup.
内容的提问来源于stack exchange,提问作者AljoSt

