You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将MKA(Matroska容器)音频文件解码为原始音频数据?

解决MKA文件提取指定大小原始音频数据的问题

你已经搞定了从MKA容器解码音频的核心逻辑,但确实如你所说,getData方法里的dataSize参数完全没派上用场——当前代码每次只返回单帧的部分数据(而且还只取了linesize[0],这对多声道或者不同采样格式的音频来说还会丢数据),没法按指定字节数返回数据。

接下来我会帮你改造代码,核心思路是维护一个内部缓存缓冲区:每次解码得到完整的音频帧数据后存入缓存,然后根据请求的大小从缓存里取出对应的数据;如果缓存不足,就继续解码新的帧填充缓存,直到满足需求或者解码结束。


步骤1:给AudioDecoder类添加缓存成员

先在audiodecoder.h的protected区域新增缓存相关的变量,用来暂存已解码的音频数据:

protected:
    bool m_initialized;
    QString m_fileName;
    AVFrame* p_frame = nullptr;
    AVPacket* p_packet = nullptr;
    AVCodecContext* p_cdcCtx = nullptr;
    AVFormatContext* p_frmCtx = nullptr;
    // 新增:用于暂存已解码音频数据的缓存
    QByteArray m_audioCache;

步骤2:修复帧数据读取的bug

原来的代码只读取了p_frame->data[0]的linesize[0]字节,这只适用于单声道平面格式的音频。对于交错格式(比如S16LE)或者多声道音频,会直接丢失大量数据。我们需要根据音频的采样格式,正确计算并读取完整的帧数据:

  • 交错格式(如AV_SAMPLE_FMT_S16):所有声道的数据都存在data[0],总大小为帧采样数 × 每个采样字节数 × 声道数
  • 平面格式(如AV_SAMPLE_FMT_FLTP):每个声道的数据存在单独的data[i],每个声道大小为linesize[0],总大小为linesize[0] × 声道数

步骤3:改造getData方法,实现按指定大小返回数据

修改后的getData会优先从缓存取数据,缓存不足时再解码新帧填充缓存,直到满足请求大小或者解码结束:

QByteArray AudioDecoder::getData(const quint16& dataSize) noexcept {
    QByteArray result;
    // 先尝试从缓存中获取指定大小的数据
    if (m_audioCache.size() >= dataSize) {
        result = m_audioCache.left(dataSize);
        m_audioCache.remove(0, dataSize);
        return result;
    }

    // 缓存不足,继续解码帧填充缓存
    while (true) {
        int response = av_read_frame(p_frmCtx, p_packet);
        if (response < 0) {
            // 解码到末尾,返回缓存中剩余的所有数据
            result = m_audioCache;
            m_audioCache.clear();
            break;
        }

        response = avcodec_send_packet(p_cdcCtx, p_packet);
        if (response < 0) {
            logging("Error while sending a packet to the decoder");
            av_packet_unref(p_packet);
            break;
        }

        while (response >= 0) {
            response = avcodec_receive_frame(p_cdcCtx, p_frame);
            if (response == AVERROR(EAGAIN) || response == AVERROR_EOF) {
                break;
            } else if (response < 0) {
                logging("Error while receiving a frame from the decoder");
                av_packet_unref(p_packet);
                return {};
            }

            // 计算当前帧的总字节数并读取完整数据
            int bytesPerSample = av_get_bytes_per_sample(p_cdcCtx->sample_fmt);
            int frameSize = 0;
            QByteArray frameData;

            if (av_sample_fmt_is_planar(p_cdcCtx->sample_fmt)) {
                // 平面格式:拼接所有声道的数据
                frameSize = p_frame->linesize[0] * p_cdcCtx->channels;
                frameData.resize(frameSize);
                uint8_t* dst = frameData.data();
                for (int ch = 0; ch < p_cdcCtx->channels; ch++) {
                    memcpy(dst, p_frame->data[ch], p_frame->linesize[0]);
                    dst += p_frame->linesize[0];
                }
            } else {
                // 交错格式:直接读取data[0]的全部数据
                frameSize = p_frame->nb_samples * bytesPerSample * p_cdcCtx->channels;
                frameData.append((char*)p_frame->data[0], frameSize);
            }

            logging(QString("Frame %1 (sample fmt=%2, channels=%3, frame size=%4 bytes) pts %5")
                    .arg(p_cdcCtx->frame_number)
                    .arg(av_get_sample_fmt_name(p_cdcCtx->sample_fmt))
                    .arg(p_cdcCtx->channels)
                    .arg(frameSize)
                    .arg(p_frame->pts)
                    .toStdString().c_str());

            // 将完整帧数据加入缓存
            m_audioCache.append(frameData);

            // 检查缓存是否满足请求大小,满足则返回数据
            if (m_audioCache.size() >= dataSize) {
                result = m_audioCache.left(dataSize);
                m_audioCache.remove(0, dataSize);
                av_packet_unref(p_packet);
                return result;
            }
        }

        av_packet_unref(p_packet);
    }

    return result;
}

步骤4:初始化缓存并修复内存泄漏

在构造函数里初始化缓存,同时完善stop方法的内存释放逻辑,避免内存泄漏:

// 构造函数初始化缓存
AudioDecoder::AudioDecoder(const QString& fileName)
    : m_initialized(false)
    , m_fileName(fileName)
    , p_cdcCtx(nullptr)
    , p_frmCtx(nullptr)
    , m_audioCache()
{
    av_register_all();
}

// 完善stop方法的内存释放
bool AudioDecoder::stop(void) noexcept {
    if (p_cdcCtx != nullptr) {
        avcodec_free_context(&p_cdcCtx);
        p_cdcCtx = nullptr;
    }
    if (p_frmCtx != nullptr) {
        avformat_close_input(&p_frmCtx);
        p_frmCtx = nullptr;
    }
    if (p_frame != nullptr) {
        av_frame_free(&p_frame);
        p_frame = nullptr;
    }
    if (p_packet != nullptr) {
        av_packet_free(&p_packet);
        p_packet = nullptr;
    }
    m_audioCache.clear();
    m_initialized = false;
    return true;
}

这样修改后,你的getData方法就会严格按照传入的size参数返回对应大小的原始音频数据,同时修复了帧数据读取不完整的问题,也优化了内存释放逻辑,减少了内存泄漏的风险。

内容的提问来源于stack exchange,提问作者bbdd

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 10:02:33