使用AVAudioEngine离线处理音频时AVAudioUnitTimePitch变速异常问题
I've run into this exact issue before—when you add time stretching/pitching to an offline AVAudioEngine render, the naive "stop when we reach the original file's length" logic breaks because the time pitch unit alters the input-to-output frame ratio. Here's why your current code fails and how to fix it completely:
Why the Original Code Breaks
- When you set
rate < 1(slow down), theAVAudioUnitTimePitchstretches the audio, meaning each input frame produces more than one output frame. Your loop stops when the input position hits the original file's length, but the time pitch unit still has buffered audio to output—so you get a truncated file. - When
rate > 1(speed up), each input frame produces fewer output frames. Your loop keeps running until the input position finishes the original file, but after that, there's no more audio to process, so you end up writing silence to the output file.
The Fix: Track Output Frames, Not Input Frames
Instead of checking against the original audio file's length, calculate how many frames you expect to output based on the original duration and your playback rate. Then run the loop until you've rendered that total number of frames.
Here's the updated version of your render(to:) function (I'll assume you have a reference to your AVAudioUnitTimePitch node as timePitchNode):
// engine: AVAudioEngine // playerNode: AVAudioPlayerNode // audioFile: AVAudioFile // timePitchNode: AVAudioUnitTimePitch (your pitch/time stretch node) open func render(to destinationFile: AVAudioFile) throws { playerNode.scheduleFile(audioFile, at: nil) do { let buffCapacity: AVAudioFrameCount = 4096 try engine.enableManualRenderingMode(.offline, format: audioFile.processingFormat, maximumFrameCount: buffCapacity) } catch { print("Failed to enable manual rendering mode: \(error)") throw error } do { try engine.start() } catch { print("Failed to start the engine: \(error)") throw error } playerNode.play() let outputBuff = AVAudioPCMBuffer(pcmFormat: engine.manualRenderingFormat, frameCapacity: engine.manualRenderingMaximumFrameCount)! let rate = timePitchNode.rate // Calculate total expected output frames based on original duration and playback rate let originalSampleRate = audioFile.processingFormat.sampleRate let originalDuration = Double(audioFile.length) / originalSampleRate let outputDuration = originalDuration / rate let totalOutputFrames = AVAudioFrameCount(outputDuration * destinationFile.processingFormat.sampleRate) var renderedFrames: AVAudioFrameCount = 0 while renderedFrames < totalOutputFrames { // Calculate how many frames we can render in this iteration (don't exceed remaining frames) let remainingOutputFrames = totalOutputFrames - renderedFrames let framesToRender = min(outputBuff.frameCapacity, remainingOutputFrames) do { let renderingStatus = try engine.renderOffline(framesToRender, to: outputBuff) switch renderingStatus { case .success: // Write only the actual frames rendered (in case we got fewer than requested) outputBuff.frameLength = framesToRender try destinationFile.write(from: outputBuff) renderedFrames += framesToRender case .insufficientDataFromInputNode: // This can happen when the input has finished but the time pitch node still has data // Try rendering smaller chunks until we get all output frames continue case .cannotDoInCurrentContext: continue case .error: print("An error occured during rendering.") throw AudioPlayer.ExportError.renderingError @unknown default: fatalError("engine.renderOffline() returned an unknown value.") } } catch { print("Failed to render offline manually: \(error)") throw error } } playerNode.stop() engine.stop() engine.disableManualRenderingMode() }
Key Changes Explained
- Calculate Expected Output Frames: We derive the total number of frames we need to render by adjusting the original audio's duration with the playback rate, then converting that back to frames using the destination file's sample rate.
- Track Rendered Output Frames: Instead of relying on
engine.manualRenderingSampleTime(which tracks input progress), we count how many frames we've actually written to the destination file. - Handle Partial Renders: The loop accounts for cases where we might get fewer frames than requested (like when the input is exhausted but the time pitch node still has buffered audio) by continuing until we hit the total output frame count.
Edge Cases to Watch For
- Make sure the
ratevalue is valid (greater than 0—AVAudioUnitTimePitch doesn't support 0 or negative rates). - If your destination file uses a different sample rate than the original, the calculation still works because we use the destination's sample rate to compute output frames.
- For very slow rates (e.g., 0.1), the time pitch node might have a larger internal buffer, so the
insufficientDataFromInputNodecase will handle flushing that buffer.
内容的提问来源于stack exchange,提问作者Fabio

