You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在AVAudioEngine中播放AAC流:转换无输出帧问题排查

问题排查:AAC转PCM时packetCount为0导致无音频输出

需要在仅支持未压缩线性PCM缓冲的AVAudioPlayerNode中播放流式AAC数据,流程为:

  1. 将AAC数据块写入AVAudioCompressedBuffer
  2. 转换为PCM缓冲
  3. 通过Player Node调度播放

但调试发现**compressedBuffer.packetCount始终为0**,最终转换后的PCM缓冲帧长度也为0,无音频输出。

实现代码

func preparePlayer() {
    guard let engine = audioEngine else {
        print("Audio engine is not initialized")
        return
    }

    // Initialize a player node and attach it to the engine
    playerNode = AVAudioPlayerNode()

    engine.attach(playerNode!)
    print("Output Node Format: \(engine.outputNode.outputFormat(forBus: 0))")

    pcmFormat = AVAudioFormat(commonFormat: .pcmFormatFloat32, sampleRate: 48000, channels: 1, interleaved: false)
    
    engine.connect(playerNode!, to: engine.outputNode, format: pcmFormat)

    var asbd = AudioStreamBasicDescription()
    asbd.mSampleRate = 24000  // 24 kHz
    asbd.mFormatID = kAudioFormatMPEG4AAC // AAC
    asbd.mChannelsPerFrame = 1  // Mono
    asbd.mBytesPerPacket = 0  // Varies (compressed format)
    asbd.mFramesPerPacket = 1024 
    asbd.mBytesPerFrame = 0  // Varies (compressed format)
    asbd.mBitsPerChannel = 0  // Varies (compressed format)

    // Create an AVAudioFormat with the AudioStreamBasicDescription
    sourceFormat = AVAudioFormat(streamDescription: &asbd)
    
    // Initialize the audio converter with the source (AAC) and destination (PCM) formats
    audioConverter = AVAudioConverter(from: sourceFormat!, to: pcmFormat!)
}

func playAudio(audioData: Data) {
    guard let playerNode = self.playerNode else {
        print("Player node is not initialized")
        return
    }
    guard let engine = audioEngine else {
        print("Audio engine is not initialized")
        return
    }
    
    let compressedBuffer = AVAudioCompressedBuffer(format: sourceFormat!, packetCapacity: 1024, maximumPacketSize: audioConverter!.maximumOutputPacketSize)
    compressedBuffer.byteLength = AVAudioPacketCount(audioData.count)
    print("audioData contains \(audioData.count) counts")
    let middleIndex = audioData.count / 2
    let middleRangeAudioData = middleIndex..<(middleIndex + 10)
    print("Middle bytes of audioData: \(Array(audioData[middleRangeAudioData]))")

    
    audioData.withUnsafeBytes {
        compressedBuffer.data.copyMemory(from: $0.baseAddress!, byteCount: audioData.count)
    }
    print("compressedBuffer contains \(compressedBuffer.packetCount) packet counts")
    print("compressedBuffer contains \(compressedBuffer.byteCapacity) byte capacity")
    print("compressedBuffer contains \(compressedBuffer.byteLength) valid bytes")
    print("compressedBuffer contains \(compressedBuffer.packetCapacity) packet capacity")
    let bufferPointer = compressedBuffer.data.bindMemory(to: UInt8.self, capacity: audioData.count)
    let bufferBytes = Array(UnsafeBufferPointer(start: bufferPointer, count: audioData.count))
    let middleRangeCompressedBuffer = middleIndex..<(middleIndex + 10)
    print("Middle bytes of compressedBuffer: \(Array(bufferBytes[middleRangeCompressedBuffer]))")


    // Create a PCM buffer
    let pcmBuffer = AVAudioPCMBuffer(pcmFormat: pcmFormat!, frameCapacity: 1024)

    let inputBlock: AVAudioConverterInputBlock = { inNumPackets, outStatus in
        outStatus.pointee = AVAudioConverterInputStatus.haveData
        return compressedBuffer
    }

    var error: NSError?
    let conversionResult = audioConverter!.convert(to: pcmBuffer!, error: &error, withInputFrom: inputBlock)
    if conversionResult == .error {
        print("Conversion failed with error: \(String(describing: error))")
    } else {
        print("Conversion successful")
        print("buffer contains \(pcmBuffer?.frameLength ?? 123456) frames")
    }
            
    if let frameLength = pcmBuffer?.frameLength, frameLength > 0 {
        let channelCount = pcmFormat?.channelCount ?? 0
        for channel in 0..<channelCount {
            if let channelData = pcmBuffer?.floatChannelData?[Int(channel)] {
                let channelDataPointer = UnsafeBufferPointer(start: channelData, count: Int(frameLength))
                let firstFewSamples = Array(channelDataPointer.prefix(10))
                print("First few samples of pcmBuffer in channel \(channel): \(firstFewSamples)")
            }
        }
    }

    
    if !engine.isRunning {
        do {
            try engine.start()
        } catch {
            print("Error starting audio engine: \(error)")
        }
    }

    playerNode.scheduleBuffer(pcmBuffer!, completionHandler: nil)
}

调试日志

audioData contains 1369 counts
Middle bytes of audioData: [250, 206, 86, 76, 254, 10, 221, 187, 190, 243]
compressedBuffer contains 0 packet counts
compressedBuffer contains 4096 byte capacity
compressedBuffer contains 1369 valid bytes
compressedBuffer contains 1024 packet capacity
Middle bytes of compressedBuffer: [250, 206, 86, 76, 254, 10, 221, 187, 190, 243]
Conversion successful
buffer contains 0 frames

核心问题分析

1. AVAudioCompressedBuffer未正确配置包信息

对于可变比特率的AAC格式,仅设置byteLength不足以让系统识别有效数据:

  • packetCount需要手动指定,系统无法自动推断AAC包数量
  • 可变长度的AAC包需要提供packetDescriptions数组,描述每个包的起始偏移和字节长度

2. AAC格式描述不完整

手动构造的AudioStreamBasicDescription缺少关键参数:

  • 未设置mFormatFlags,无法指定AAC的具体编码类型(如LC、HE-AAC)
  • 流式AAC通常为ADTS格式,需要格式描述正确兼容ADTS头解析

3. 转换输入块逻辑错误

输入块每次返回同一个完整缓冲,未响应转换器的inNumPackets请求,也未标记数据耗尽状态,导致转换器无法正确处理数据。


修复方案

步骤1:完善AAC格式初始化

添加AAC编码类型的格式标志,确保格式描述正确:

var asbd = AudioStreamBasicDescription()
asbd.mSampleRate = 24000
asbd.mFormatID = kAudioFormatMPEG4AAC
asbd.mChannelsPerFrame = 1
asbd.mBytesPerPacket = 0
asbd.mFramesPerPacket = 1024
asbd.mBytesPerFrame = 0
asbd.mBitsPerChannel = 0
// 指定AAC LC编码(根据实际AAC类型调整)
asbd.mFormatFlags = kMPEG4Object_AAC_LC
sourceFormat = AVAudioFormat(streamDescription: &asbd)

步骤2:解析ADTS包并配置AVAudioCompressedBuffer

流式AAC多为ADTS格式,需解析ADTS头获取包数量和每个包的长度:

// 解析ADTS格式的AAC包数量及每个包的信息
func parseADTSPackets(from data: Data) -> (packetCount: AVAudioPacketCount, descriptions: [AudioStreamPacketDescription]) {
    var packetCount: AVAudioPacketCount = 0
    var descriptions = [AudioStreamPacketDescription]()
    var offset = 0
    let dataCount = data.count
    
    while offset <= dataCount - 7 {
        // 检查ADTS同步字(0xFFF0)
        let syncWord = (UInt16(data[offset]) << 8) | UInt16(data[offset+1])
        guard syncWord & 0xFFF0 == 0xFFF0 else {
            offset += 1
            continue
        }
        
        // 解析包长度
        let lengthByte1 = data[offset+3]
        let lengthByte2 = data[offset+4]
        let packetLength = ((UInt16(lengthByte1) & 0x03) << 8) | UInt16(lengthByte2)
        
        // 添加包描述
        var desc = AudioStreamPacketDescription()
        desc.mStartOffset = UInt32(offset)
        desc.mDataByteSize = packetLength
        desc.mVariableFramesInPacket = 0
        descriptions.append(desc)
        
        packetCount += 1
        offset += Int(packetLength)
    }
    return (packetCount, descriptions)
}

在playAudio中应用解析结果:

let (packetCount, packetDescriptions) = parseADTSPackets(from: audioData)
compressedBuffer.packetCount = packetCount
// 写入包描述(仅当包长度可变时需要)
if !packetDescriptions.isEmpty {
    compressedBuffer.packetDescriptions = UnsafeMutablePointer<AudioStreamPacketDescription>.allocate(capacity: packetDescriptions.count)
    for i in 0..<packetDescriptions.count {
        compressedBuffer.packetDescriptions[i] = packetDescriptions[i]
    }
}

步骤3:修正转换输入块逻辑

让输入块响应转换器的请求,逐步提供数据:

var currentPacketIndex: AVAudioPacketCount = 0
let inputBlock: AVAudioConverterInputBlock = { inNumPackets, outStatus in
    guard currentPacketIndex < packetCount else {
        outStatus.pointee = .noDataNow
        return nil
    }
    
    let packetsToProvide = min(AVAudioPacketCount(inNumPackets), packetCount - currentPacketIndex)
    // 创建子缓冲提供对应数量的包
    let subBuffer = AVAudioCompressedBuffer(format: sourceFormat!, packetCapacity: packetsToProvide, maximumPacketSize: compressedBuffer.maximumPacketSize)
    
    // 复制对应的数据段
    let startOffset = Int(compressedBuffer.packetDescriptions[Int(currentPacketIndex)].mStartOffset)
    let endOffset = startOffset + Int(compressedBuffer.packetDescriptions[Int(currentPacketIndex + packetsToProvide - 1)].mStartOffset + compressedBuffer.packetDescriptions[Int(currentPacketIndex + packetsToProvide - 1)].mDataByteSize)
    let subData = audioData.subdata(in: startOffset..<endOffset)
    subBuffer.byteLength = AVAudioPacketCount(subData.count)
    subBuffer.packetCount = packetsToProvide
    
    subData.withUnsafeBytes {
        subBuffer.data.copyMemory(from: $0.baseAddress!, byteCount: subData.count)
    }
    
    currentPacketIndex += packetsToProvide
    outStatus.pointee = .haveData
    return subBuffer
}

步骤4:调整PCM缓冲容量

AAC每包对应1024帧,设置足够的PCM缓冲容量:

let pcmBuffer = AVAudioPCMBuffer(pcmFormat: pcmFormat!, frameCapacity: AVAudioFrameCount(packetCount * 1024))

关键注意点

  • 流式AAC必须解析ADTS头才能获取包信息,不能直接写入原始数据
  • AVAudioConverter依赖明确的包数量和包描述才能完成解码
  • 调试时可打印sourceFormat.debugDescription确认格式参数是否正确

内容的提问来源于stack exchange,提问作者Bardigan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 19:24:56