You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

WebRTC I420Frame转YUV_420_888/NV21及帧数据获取问题

问题描述

我基于AndroidWebRTC项目,使用libjingle_peerconnection.jar实现了视频通话,运行一切正常。现在需要实时检测人脸及微笑,打算用ML Kit完成此功能,因此需要获取每帧的ByteBuffer。

Google官方说明:

若使用Camera2 API,请以ImageFormat.YUV_420_888格式捕获图像;若使用旧版Camera API,请以ImageFormat.NV21格式捕获图像。

为获取帧数据,我把SurfaceViewRenderer的源码复制到项目中使用,该类包含一个接收org.webrtc.VideoRenderer.I420Frame作为参数的renderFrame方法。目前遇到两个问题:

  • 使用本地摄像头时,帧的yuvPlanes值为null;
  • 接收远程视频时,yuvPlanes值不为null,但无法将其转换为YUV_420_888和NV21格式,请问该如何处理?

I420Frame类属性如下:

public static class I420Frame {
    public final int width;
    public final int height;
    public final int[] yuvStrides;
    public ByteBuffer[] yuvPlanes;
    public final boolean yuvFrame;
    public final float[] samplingMatrix;
    public int textureId;
    private long nativeFramePointer;
    public int rotationDegree;
}
解决方案

问题1:本地摄像头帧yuvPlanes为null

WebRTC默认用纹理渲染本地摄像头画面,此时帧数据以textureId传递,不会填充yuvPlanes。解决方式有两种:

方式1:强制WebRTC输出I420格式帧

在初始化本地视频捕获器时,添加约束强制使用I420编码格式:

// 以Camera1为例,创建捕获器并设置约束
CameraVideoCapturer capturer = new Camera1Enumerator(true).createCapturer("0", null);
MediaConstraints constraints = new MediaConstraints();
// 基础分辨率、帧率约束
constraints.mandatory.add(new MediaConstraints.KeyValuePair("maxWidth", "640"));
constraints.mandatory.add(new MediaConstraints.KeyValuePair("maxHeight", "480"));
constraints.mandatory.add(new MediaConstraints.KeyValuePair("maxFrameRate", "30"));
// 关键:强制输出I420格式
constraints.mandatory.add(new MediaConstraints.KeyValuePair("videoCodec", "I420"));

// 初始化VideoSource并启动捕获
VideoSource videoSource = peerConnectionFactory.createVideoSource(capturer.isScreencast());
videoSource.adaptOutputFormat(640, 480, 30);
capturer.startCapture(640, 480, 30, constraints);

配置后,renderFrame接收的I420Frame会自动填充yuvPlanes数据。

方式2:将纹理帧转换为I420格式

如果无法修改捕获配置,可以通过WebRTC的native指针将纹理帧转为I420 ByteBuffer:

  1. 编写JNI方法提取YUV数据:
extern "C" JNIEXPORT void JNICALL
Java_com_your_package_YourSurfaceViewRenderer_convertTextureToI420(
        JNIEnv* env, jobject thiz, jlong native_frame,
        jobject y_buffer, jobject u_buffer, jobject v_buffer, jintArray strides) {
    webrtc::VideoFrame* frame = reinterpret_cast<webrtc::VideoFrame*>(native_frame);
    const webrtc::I420BufferInterface* i420_buffer = frame->video_frame_buffer()->ToI420();

    // 复制Y平面
    uint8_t* y_data = static_cast<uint8_t*>(env->GetDirectBufferAddress(y_buffer));
    memcpy(y_data, i420_buffer->DataY(), i420_buffer->StrideY() * i420_buffer->height());

    // 复制U平面
    uint8_t* u_data = static_cast<uint8_t*>(env->GetDirectBufferAddress(u_buffer));
    memcpy(u_data, i420_buffer->DataU(), i420_buffer->StrideU() * (i420_buffer->height() / 2));

    // 复制V平面
    uint8_t* v_data = static_cast<uint8_t*>(env->GetDirectBufferAddress(v_buffer));
    memcpy(v_data, i420_buffer->DataV(), i420_buffer->StrideV() * (i420_buffer->height() / 2));

    // 设置strides
    jint* stride_arr = env->GetIntArrayElements(strides, nullptr);
    stride_arr[0] = i420_buffer->StrideY();
    stride_arr[1] = i420_buffer->StrideU();
    stride_arr[2] = i420_buffer->StrideV();
    env->ReleaseIntArrayElements(strides, stride_arr, 0);
}
  1. 在SurfaceViewRenderer的renderFrame方法中调用转换:
private native void convertTextureToI420(long nativeFrame, ByteBuffer yPlane, ByteBuffer uPlane, ByteBuffer vPlane, int[] strides);

@Override
public void renderFrame(I420Frame frame) {
    // 处理纹理转I420逻辑
    if (frame.yuvPlanes == null && frame.nativeFramePointer != 0) {
        int ySize = frame.width * frame.height;
        int uvSize = ySize / 4;
        ByteBuffer[] yuvPlanes = new ByteBuffer[3];
        yuvPlanes[0] = ByteBuffer.allocateDirect(ySize);
        yuvPlanes[1] = ByteBuffer.allocateDirect(uvSize);
        yuvPlanes[2] = ByteBuffer.allocateDirect(uvSize);
        int[] strides = new int[3];
        
        convertTextureToI420(frame.nativeFramePointer, yuvPlanes[0], yuvPlanes[1], yuvPlanes[2], strides);
        frame.yuvPlanes = yuvPlanes;
        frame.yuvStrides = strides;
    }

    // 原有渲染逻辑...

    // 此处可使用yuvPlanes进行ML Kit检测
}

问题2:远程视频帧转换为ML Kit支持的格式

远程视频的yuvPlanes是标准I420格式(Y、U、V三个独立平面),可直接转换为目标格式:

转换为YUV_420_888

直接用ML Kit的InputImage.fromByteBuffer传入I420数据:

// 重置缓冲区指针
frame.yuvPlanes[0].rewind();
frame.yuvPlanes[1].rewind();
frame.yuvPlanes[2].rewind();

// 创建InputImage
InputImage inputImage = InputImage.fromByteBuffer(
        frame.yuvPlanes[0],
        frame.width,
        frame.height,
        frame.rotationDegree,
        InputImage.IMAGE_FORMAT_I420
);

// ML Kit人脸检测
FaceDetector detector = FaceDetection.getClient();
detector.process(inputImage)
        .addOnSuccessListener(faces -> {
            for (Face face : faces) {
                Smile smile = face.getSmilingProbability();
                if (smile != null && smile.getValue() > 0.7) {
                    // 检测到微笑
                }
            }
        })
        .addOnFailureListener(e -> {
            // 处理错误
        });

转换为NV21格式

NV21是Y平面在前、UV交错在后的格式,需合并U、V平面:

int ySize = frame.width * frame.height;
int uvSize = ySize / 4;
ByteBuffer nv21Buffer = ByteBuffer.allocateDirect(ySize + uvSize * 2);

// 复制Y平面
frame.yuvPlanes[0].rewind();
nv21Buffer.put(frame.yuvPlanes[0]);

// 合并U、V为UV交错格式
frame.yuvPlanes[1].rewind();
frame.yuvPlanes[2].rewind();
byte[] uBytes = new byte[uvSize];
byte[] vBytes = new byte[uvSize];
frame.yuvPlanes[1].get(uBytes);
frame.yuvPlanes[2].get(vBytes);

for (int i = 0; i < uvSize; i++) {
    nv21Buffer.put(vBytes[i]);
    nv21Buffer.put(uBytes[i]);
}
nv21Buffer.rewind();

// 创建InputImage
InputImage inputImage = InputImage.fromByteBuffer(
        nv21Buffer,
        frame.width,
        frame.height,
        frame.rotationDegree,
        InputImage.IMAGE_FORMAT_NV21
);

// 后续ML Kit检测逻辑同上

内容的提问来源于stack exchange,提问作者Hussein Yaqoobi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 04:37:15