如何将MKA(Matroska容器)音频文件解码为原始音频数据?
解决MKA文件提取指定大小原始音频数据的问题
你已经搞定了从MKA容器解码音频的核心逻辑,但确实如你所说,getData方法里的dataSize参数完全没派上用场——当前代码每次只返回单帧的部分数据(而且还只取了linesize[0],这对多声道或者不同采样格式的音频来说还会丢数据),没法按指定字节数返回数据。
接下来我会帮你改造代码,核心思路是维护一个内部缓存缓冲区:每次解码得到完整的音频帧数据后存入缓存,然后根据请求的大小从缓存里取出对应的数据;如果缓存不足,就继续解码新的帧填充缓存,直到满足需求或者解码结束。
步骤1:给AudioDecoder类添加缓存成员
先在audiodecoder.h的protected区域新增缓存相关的变量,用来暂存已解码的音频数据:
protected: bool m_initialized; QString m_fileName; AVFrame* p_frame = nullptr; AVPacket* p_packet = nullptr; AVCodecContext* p_cdcCtx = nullptr; AVFormatContext* p_frmCtx = nullptr; // 新增:用于暂存已解码音频数据的缓存 QByteArray m_audioCache;
步骤2:修复帧数据读取的bug
原来的代码只读取了p_frame->data[0]的linesize[0]字节,这只适用于单声道平面格式的音频。对于交错格式(比如S16LE)或者多声道音频,会直接丢失大量数据。我们需要根据音频的采样格式,正确计算并读取完整的帧数据:
- 交错格式(如AV_SAMPLE_FMT_S16):所有声道的数据都存在
data[0],总大小为帧采样数 × 每个采样字节数 × 声道数 - 平面格式(如AV_SAMPLE_FMT_FLTP):每个声道的数据存在单独的
data[i],每个声道大小为linesize[0],总大小为linesize[0] × 声道数
步骤3:改造getData方法,实现按指定大小返回数据
修改后的getData会优先从缓存取数据,缓存不足时再解码新帧填充缓存,直到满足请求大小或者解码结束:
QByteArray AudioDecoder::getData(const quint16& dataSize) noexcept { QByteArray result; // 先尝试从缓存中获取指定大小的数据 if (m_audioCache.size() >= dataSize) { result = m_audioCache.left(dataSize); m_audioCache.remove(0, dataSize); return result; } // 缓存不足,继续解码帧填充缓存 while (true) { int response = av_read_frame(p_frmCtx, p_packet); if (response < 0) { // 解码到末尾,返回缓存中剩余的所有数据 result = m_audioCache; m_audioCache.clear(); break; } response = avcodec_send_packet(p_cdcCtx, p_packet); if (response < 0) { logging("Error while sending a packet to the decoder"); av_packet_unref(p_packet); break; } while (response >= 0) { response = avcodec_receive_frame(p_cdcCtx, p_frame); if (response == AVERROR(EAGAIN) || response == AVERROR_EOF) { break; } else if (response < 0) { logging("Error while receiving a frame from the decoder"); av_packet_unref(p_packet); return {}; } // 计算当前帧的总字节数并读取完整数据 int bytesPerSample = av_get_bytes_per_sample(p_cdcCtx->sample_fmt); int frameSize = 0; QByteArray frameData; if (av_sample_fmt_is_planar(p_cdcCtx->sample_fmt)) { // 平面格式:拼接所有声道的数据 frameSize = p_frame->linesize[0] * p_cdcCtx->channels; frameData.resize(frameSize); uint8_t* dst = frameData.data(); for (int ch = 0; ch < p_cdcCtx->channels; ch++) { memcpy(dst, p_frame->data[ch], p_frame->linesize[0]); dst += p_frame->linesize[0]; } } else { // 交错格式:直接读取data[0]的全部数据 frameSize = p_frame->nb_samples * bytesPerSample * p_cdcCtx->channels; frameData.append((char*)p_frame->data[0], frameSize); } logging(QString("Frame %1 (sample fmt=%2, channels=%3, frame size=%4 bytes) pts %5") .arg(p_cdcCtx->frame_number) .arg(av_get_sample_fmt_name(p_cdcCtx->sample_fmt)) .arg(p_cdcCtx->channels) .arg(frameSize) .arg(p_frame->pts) .toStdString().c_str()); // 将完整帧数据加入缓存 m_audioCache.append(frameData); // 检查缓存是否满足请求大小,满足则返回数据 if (m_audioCache.size() >= dataSize) { result = m_audioCache.left(dataSize); m_audioCache.remove(0, dataSize); av_packet_unref(p_packet); return result; } } av_packet_unref(p_packet); } return result; }
步骤4:初始化缓存并修复内存泄漏
在构造函数里初始化缓存,同时完善stop方法的内存释放逻辑,避免内存泄漏:
// 构造函数初始化缓存 AudioDecoder::AudioDecoder(const QString& fileName) : m_initialized(false) , m_fileName(fileName) , p_cdcCtx(nullptr) , p_frmCtx(nullptr) , m_audioCache() { av_register_all(); } // 完善stop方法的内存释放 bool AudioDecoder::stop(void) noexcept { if (p_cdcCtx != nullptr) { avcodec_free_context(&p_cdcCtx); p_cdcCtx = nullptr; } if (p_frmCtx != nullptr) { avformat_close_input(&p_frmCtx); p_frmCtx = nullptr; } if (p_frame != nullptr) { av_frame_free(&p_frame); p_frame = nullptr; } if (p_packet != nullptr) { av_packet_free(&p_packet); p_packet = nullptr; } m_audioCache.clear(); m_initialized = false; return true; }
这样修改后,你的getData方法就会严格按照传入的size参数返回对应大小的原始音频数据,同时修复了帧数据读取不完整的问题,也优化了内存释放逻辑,减少了内存泄漏的风险。
内容的提问来源于stack exchange,提问作者bbdd
相关产品推荐
相关产品推荐

