You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从Y'Cb'Cr编码视频快速提取I帧原始Y'分量数据?

解决方案:直接提取视频I帧原始Y'分量到内存生成感知哈希

核心思路

  • 跳过YUV→RGB转换和磁盘中间文件,直接从解码后的帧中提取原始Y'平面数据
  • 利用FFmpeg的libav*系列库在C++中实现,直接将Y'数据写入内存数组
  • 关联每个I帧的帧号,生成哈希后立即释放内存,避免占用过多资源

具体实现步骤

1. 初始化FFmpeg上下文

先初始化FFmpeg核心组件,打开视频文件并定位视频流:

#include <iostream>
#include <vector>
#include <cstdint>
extern "C" {
#include <libavformat/avformat.h>
#include <libavcodec/avcodec.h>
#include <libavutil/imgutils.h>
}

int main() {
    av_register_all();
    avformat_network_init();

    AVFormatContext* fmt_ctx = nullptr;
    if (avformat_open_input(&fmt_ctx, "input.mp4", nullptr, nullptr) != 0) {
        std::cerr << "Failed to open video file" << std::endl;
        return -1;
    }

    if (avformat_find_stream_info(fmt_ctx, nullptr) < 0) {
        std::cerr << "Failed to get stream info" << std::endl;
        avformat_close_input(&fmt_ctx);
        return -1;
    }

    // 定位视频流索引
    int video_stream_idx = -1;
    for (int i = 0; i < fmt_ctx->nb_streams; ++i) {
        if (fmt_ctx->streams[i]->codecpar->codec_type == AVMEDIA_TYPE_VIDEO) {
            video_stream_idx = i;
            break;
        }
    }
    if (video_stream_idx == -1) {
        std::cerr << "No video stream found" << std::endl;
        avformat_close_input(&fmt_ctx);
        return -1;
    }

2. 初始化解码器

读取视频流的解码器参数,打开对应解码器:

AVCodecParameters* codec_par = fmt_ctx->streams[video_stream_idx]->codecpar;
    const AVCodec* codec = avcodec_find_decoder(codec_par->codec_id);
    if (!codec) {
        std::cerr << "Unsupported codec" << std::endl;
        avformat_close_input(&fmt_ctx);
        return -1;
    }

    AVCodecContext* codec_ctx = avcodec_alloc_context3(codec);
    if (!codec_ctx) {
        std::cerr << "Failed to allocate codec context" << std::endl;
        avformat_close_input(&fmt_ctx);
        return -1;
    }

    if (avcodec_parameters_to_context(codec_ctx, codec_par) < 0) {
        std::cerr << "Failed to copy codec parameters" << std::endl;
        avcodec_free_context(&codec_ctx);
        avformat_close_input(&fmt_ctx);
        return -1;
    }

    if (avcodec_open2(codec_ctx, codec, nullptr) < 0) {
        std::cerr << "Failed to open codec" << std::endl;
        avcodec_free_context(&codec_ctx);
        avformat_close_input(&fmt_ctx);
        return -1;
    }

3. 读取解码帧,提取I帧Y'分量

遍历视频数据包,解码后筛选I帧,直接提取原始Y'平面数据到内存:

AVPacket pkt;
    AVFrame* frame = av_frame_alloc();
    while (av_read_frame(fmt_ctx, &pkt) >= 0) {
        if (pkt.stream_index != video_stream_idx) {
            av_packet_unref(&pkt);
            continue;
        }

        int ret = avcodec_send_packet(codec_ctx, &pkt);
        if (ret < 0) {
            std::cerr << "Failed to send packet for decoding" << std::endl;
            av_packet_unref(&pkt);
            break;
        }

        while (ret >= 0) {
            ret = avcodec_receive_frame(codec_ctx, frame);
            if (ret == AVERROR(EAGAIN) || ret == AVERROR_EOF) {
                break;
            } else if (ret < 0) {
                std::cerr << "Failed to receive frame" << std::endl;
                goto cleanup;
            }

            // 判断是否为I帧
            if (frame->pict_type == AV_PICTURE_TYPE_I) {
                // YUV420p格式下,frame->data[0]即为原始Y'分量
                int width = frame->width;
                int height = frame->height;
                int frame_num = frame->coded_picture_number; // 获取对应帧号

                // 复制Y'数据到内存数组,跳过FFmpeg的行对齐填充
                std::vector<uint8_t> y_data(height * width);
                for (int y = 0; y < height; ++y) {
                    memcpy(&y_data[y * width], &frame->data[0][y * frame->linesize[0]], width);
                }

                // 此处调用你的感知哈希生成函数,传入y_data、宽高和帧号
                // generate_perceptual_hash(y_data, width, height, frame_num);

                // 哈希生成后,vector自动释放内存,无需手动处理
            }

            av_frame_unref(frame);
        }

        av_packet_unref(&pkt);
    }

cleanup:
    av_frame_free(&frame);
    avcodec_free_context(&codec_ctx);
    avformat_close_input(&fmt_ctx);
    avformat_network_deinit();
    return 0;
}

关键注意事项

  • 原始Y'数据提取:YUV420p格式下,frame->data[0]是未经过任何转换的原始Y'平面,直接复制即可,完全避免滤镜带来的额外计算
  • 行对齐处理:frame->linesize[0]可能因内存对齐大于实际帧宽度,复制时需按实际宽度逐行拷贝,避免包含填充数据
  • 内存效率:用std::vector存储Y'数据,生成哈希后自动销毁,适合批量处理大量视频时的内存管控
  • 帧号关联:通过frame->coded_picture_number或frame->pts获取帧序号,保证哈希与对应帧的关联准确性

编译说明

编译时需链接FFmpeg相关库,示例命令:

g++ -o extract_i_frame_y extract_i_frame_y.cpp -lavformat -lavcodec -lavutil

内容的提问来源于stack exchange,提问作者memeko

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 11:43:29