音频重采样缩放与WebM+Opus复用技术问题求助
FFmpeg 音频处理两大问题解决方案
一、音频重采样失真卡顿问题
问题分析
你遇到的核心问题是重采样输出的帧大小不符合Opus编码器的预期规范:
- 当你用
av_rescale_rnd(input_frame->nb_samples + swr_get_delay(resampler_context, 41100), 48000, 41100, AV_ROUND_UP)计算得到1196采样时,这个非标准帧长(Opus在48kHz下常用固定帧长如960/1920采样,对应20/40ms)会导致编码器无法正确处理,进而引发时间戳混乱、音频减速失真。 - 直接用
input_frame->nb_samples = 1024时,虽然播放暂时正常,但这是编码器内部缓存适配的“权宜之计”,并非规范做法,后续可能出现其他兼容性问题。
另外,swr_get_delay的使用也存在误区:它返回的延迟采样数是基于输入采样率的,你直接加到输入帧采样数再转输出采样率,会导致每次输出帧长波动,进一步加剧编码器的处理混乱。
解决方案
规范的做法是让重采样输出的帧长严格符合Opus的固定帧长要求(比如20ms对应960采样@48kHz),具体实现步骤:
- 预先计算输入采样率转输出采样率的比例,缓存足够的输入采样后再进行重采样
- 确保重采样输出为固定的标准帧长,避免编码器时序混乱
修改后的核心代码示例:
// 定义Opus标准帧长(20ms @48kHz) #define OPUS_STANDARD_FRAME_SIZE 960 // 计算需要的输入采样数:将输出帧长转换为输入采样率的对应值 int required_input_samples = av_rescale_rnd(OPUS_STANDARD_FRAME_SIZE, 41100, 48000, AV_ROUND_UP); // 初始化输入采样缓存,用于凑够足够采样再重采样 static AVFrame *input_cache = NULL; if (!input_cache) { input_cache = av_frame_alloc(); input_cache->format = input_frame->format; input_cache->sample_rate = input_frame->sample_rate; input_cache->channels = input_frame->channels; av_frame_get_buffer(input_cache, 0); } // 将当前输入帧数据拷贝到缓存 av_samples_copy(input_cache->data, input_frame->data, 0, input_cache->nb_samples, input_frame->nb_samples, input_frame->channels, input_frame->format); input_cache->nb_samples += input_frame->nb_samples; // 当缓存采样数达标时,执行重采样 if (input_cache->nb_samples >= required_input_samples) { AVFrame *output_frame = av_frame_alloc(); output_frame->format = output_format; output_frame->sample_rate = 48000; output_frame->channels = input_frame->channels; av_frame_get_buffer(output_frame, 0); // 执行重采样,输出固定帧长 int ret = swr_convert(swr_ctx, output_frame->data, OPUS_STANDARD_FRAME_SIZE, (const uint8_t **)input_cache->data, required_input_samples); if (ret >= 0) { output_frame->nb_samples = ret; // 将输出帧送入编码器处理 encode_audio(output_frame); } // 更新缓存:移除已处理的采样,保留剩余部分 int remaining_samples = input_cache->nb_samples - required_input_samples; memmove(input_cache->data[0], input_cache->data[0] + required_input_samples * av_get_bytes_per_sample(input_frame->format), remaining_samples * av_get_bytes_per_sample(input_frame->format)); input_cache->nb_samples = remaining_samples; av_frame_free(&output_frame); }
二、WebM+Opus复用时长显示与编码警告问题
问题分析
这个问题本质是容器时间码标准与编码器时间码标准不匹配:
- WebM容器强制要求时间码以毫秒为单位(即
time_base = 1/1000),你之前设置AVStream->time_base = 1/48000,播放器会错误地将采样率单位的时间码解析为毫秒,导致时长显示严重失真(比如17秒被误算为16分11秒)。 - 当你直接把
pts加20/30时,虽然符合WebM的毫秒时间码,但Opus编码器期望的是基于采样率的pts(每帧960采样对应pts+960),这就导致编码器接收到的时序混乱,触发Queue input is backward in time警告。
解决方案
需要将编码器的时间码与容器的时间码做规范转换,核心步骤:
- 给WebM的
AVStream设置正确的time_base = 1/1000 - Opus编码器的
codec_ctx->time_base保持为1/48000(符合编码器自身时序要求) - 编码得到
AVPacket后,将其时间戳从编码器的time_base转换为容器的time_base
修改后的关键代码示例:
// 初始化WebM音频流时,设置符合容器要求的time_base AVStream *audio_stream = avformat_new_stream(ofmt_ctx, NULL); audio_stream->time_base = (AVRational){1, 1000}; // WebM要求毫秒单位时间码 // 初始化Opus编码器时,设置符合编码器要求的time_base AVCodecContext *codec_ctx = avcodec_alloc_context3(codec); codec_ctx->time_base = (AVRational){1, 48000}; // Opus基于采样率的时序 codec_ctx->sample_rate = 48000; codec_ctx->channels = 2; codec_ctx->channel_layout = AV_CH_LAYOUT_STEREO; codec_ctx->bit_rate = 128000; codec_ctx->codec_id = AV_CODEC_ID_OPUS; // 编码AVFrame时,pts基于编码器的time_base(每帧960采样,pts累加960) static int64_t current_pts = 0; frame->pts = current_pts; current_pts += OPUS_STANDARD_FRAME_SIZE; // 每帧对应960采样 // 编码得到AVPacket后,转换时间戳到容器的time_base AVPacket pkt = {0}; av_init_packet(&pkt); int ret = avcodec_send_frame(codec_ctx, frame); if (ret >= 0) { ret = avcodec_receive_packet(codec_ctx, &pkt); if (ret >= 0) { // 将数据包的时间戳从编码器基准转换为容器基准 av_packet_rescale_ts(&pkt, codec_ctx->time_base, audio_stream->time_base); pkt.stream_index = audio_stream->index; // 写入WebM容器 av_interleaved_write_frame(ofmt_ctx, &pkt); av_packet_unref(&pkt); } }
这样既满足了Opus编码器对时序的要求,又符合WebM容器的时间码标准,时长显示正常,也不会再出现时序警告。
内容的提问来源于stack exchange,提问作者siods333333
相关产品推荐
相关产品推荐

