You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于FFmpeg libav*/libswresample的音频转码重采样故障求助

音频转码+重采样卡顿问题排查与参考方案

常见卡顿原因排查

以下是你可能遗漏的关键环节:

1. SwrContext初始化参数不完整

仅设置采样率远远不够,必须明确指定输入/输出的声道布局、样本格式,否则SwrContext会使用默认值,导致重采样计算错误。比如输入是立体声(AV_CH_LAYOUT_STEREO)、样本格式为AV_SAMPLE_FMT_S16,输出是同样声道但采样率22050、样本格式为AV_SAMPLE_FMT_FLTP(AAC编码器常用格式),必须在初始化时全部指定:

SwrContext *swr_ctx = swr_alloc_set_opts(NULL,
    out_codec_ctx->channel_layout, out_codec_ctx->sample_fmt, out_codec_ctx->sample_rate,
    in_codec_ctx->channel_layout, in_codec_ctx->sample_fmt, in_codec_ctx->sample_rate,
    0, NULL);
if (!swr_ctx || swr_init(swr_ctx) < 0) {
    // 错误处理
}

2. 样本数计算逻辑错误

计算输出样本数时,必须结合输入帧的样本数和SwrContext的延迟,正确的公式应该是:

int64_t delay = swr_get_delay(swr_ctx, in_codec_ctx->sample_rate);
int out_nb_samples = av_rescale_rnd(delay + in_frame->nb_samples, out_codec_ctx->sample_rate, in_codec_ctx->sample_rate, AV_ROUND_UP);

注意这里用swr_get_delay而非直接swr_delay()(后者返回的延迟样本数需结合采样率转换),并且用AV_ROUND_UP确保不会截断样本。

3. 时间戳(PTS/DTS)未正确转换

重采样后的帧必须调整时间戳,否则播放器会因时序混乱导致卡顿。转换公式:

out_frame->pts = av_rescale_q(in_frame->pts, in_codec_ctx->time_base, out_codec_ctx->time_base);
// 基于样本数的更准确计算方式:
out_frame->pts = av_rescale_q(swr_get_delay(swr_ctx, in_codec_ctx->sample_rate) + in_frame->pts, 
                              av_make_q(1, in_codec_ctx->sample_rate), 
                              out_codec_ctx->time_base);

4. 未处理Flush阶段的剩余样本

转码结束时,必须分别Flush解码器、重采样器、编码器,获取所有剩余样本:

  • 解码Flush:向解码器输入NULL包,获取剩余解码帧
  • 重采样Flush:调用swr_convert(swr_ctx, out_data, out_nb_samples, NULL, 0)获取剩余重采样样本
  • 编码Flush:向编码器输入NULL帧,获取剩余编码包

完整转码+重采样核心流程

// 初始化输入、输出、解码器、编码器、SwrContext(略)

AVFrame *in_frame = av_frame_alloc();
AVFrame *out_frame = av_frame_alloc();
AVPacket in_pkt = {0}, out_pkt = {0};

while (av_read_frame(in_ctx, &in_pkt) >= 0) {
    // 解码输入包
    if (avcodec_send_packet(in_codec_ctx, &in_pkt) >= 0) {
        while (avcodec_receive_frame(in_codec_ctx, in_frame) >= 0) {
            // 计算输出样本数
            int64_t delay = swr_get_delay(swr_ctx, in_codec_ctx->sample_rate);
            int out_nb_samples = av_rescale_rnd(delay + in_frame->nb_samples, 
                                                out_codec_ctx->sample_rate, 
                                                in_codec_ctx->sample_rate, 
                                                AV_ROUND_UP);

            // 分配输出帧缓冲区
            out_frame->nb_samples = out_nb_samples;
            out_frame->format = out_codec_ctx->sample_fmt;
            out_frame->channel_layout = out_codec_ctx->channel_layout;
            if (av_frame_get_buffer(out_frame, 0) < 0) {
                // 错误处理
            }

            // 执行重采样
            int converted_samples = swr_convert(swr_ctx, 
                                               out_frame->data, out_frame->nb_samples,
                                               (const uint8_t **)in_frame->data, in_frame->nb_samples);
            if (converted_samples < 0) {
                // 错误处理
            }

            // 调整时间戳
            out_frame->pts = av_rescale_q(in_frame->pts, in_codec_ctx->time_base, out_codec_ctx->time_base);

            // 编码输出帧
            if (avcodec_send_frame(out_codec_ctx, out_frame) >= 0) {
                while (avcodec_receive_packet(out_codec_ctx, &out_pkt) >= 0) {
                    // 调整输出包时间戳
                    av_packet_rescale_ts(&out_pkt, out_codec_ctx->time_base, out_ctx->streams[0]->time_base);
                    out_pkt.stream_index = 0;
                    // 写入输出文件
                    av_interleaved_write_frame(out_ctx, &out_pkt);
                    av_packet_unref(&out_pkt);
                }
            }
            av_frame_unref(out_frame);
        }
    }
    av_packet_unref(&in_pkt);
}

// Flush阶段:处理剩余数据
// Flush解码器
avcodec_send_packet(in_codec_ctx, NULL);
while (avcodec_receive_frame(in_codec_ctx, in_frame) >= 0) {
    // 重复上述重采样、编码流程
}

// Flush重采样器
int out_nb_samples = swr_get_delay(swr_ctx, in_codec_ctx->sample_rate);
out_nb_samples = av_rescale_rnd(out_nb_samples, out_codec_ctx->sample_rate, in_codec_ctx->sample_rate, AV_ROUND_UP);
out_frame->nb_samples = out_nb_samples;
av_frame_get_buffer(out_frame, 0);
swr_convert(swr_ctx, out_frame->data, out_frame->nb_samples, NULL, 0);
// 编码该帧(略)

// Flush编码器
avcodec_send_frame(out_codec_ctx, NULL);
while (avcodec_receive_packet(out_codec_ctx, &out_pkt) >= 0) {
    // 写入输出文件(略)
}

// 资源释放(略)

额外注意事项

  • 确保输入输出的声道数匹配,若需要声道数转换(比如立体声转单声道),需在SwrContext初始化时指定正确的声道布局。
  • 样本格式必须与编码器要求一致,比如AAC编码器通常要求AV_SAMPLE_FMT_FLTP(平面浮点格式),如果输入是AV_SAMPLE_FMT_S16,SwrContext会自动转换,但必须明确指定。
  • 避免在每次转换时重新分配SwrContext,应初始化一次后复用。

内容的提问来源于stack exchange,提问作者whatdoido

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 12:06:29