Google Meet测试版背景模糊功能技术实现探究及同类模型问询
Great question! Your observation about Google Meet's beta background blur feature relying on WASM-powered mediapipe_wasm_simd.wasm to call TensorFlow Lite models (segm_heavy.tflite, segm_lite.tflite) and running on HTMLCanvasElement is totally accurate—this on-device, browser-based computer vision pipeline is standard for real-time privacy-focused features like background blur.
As you noted in your update, BodyPix is an excellent match for replicating this effect. Built on TensorFlow.js, it’s purpose-built for real-time person segmentation in the browser. The backgroundBlurAmount control in its demo lets you dial in the exact blur intensity to match Meet’s official behavior, and since it runs directly in the browser (with optimized variants for speed/accuracy), it mirrors Meet’s local processing flow perfectly.
Beyond BodyPix, here are other relevant tools and models that follow the same pattern:
- MediaPipe Selfie Segmentation: This is actually the foundational tech behind Google Meet’s background effects—those
segm_*.tflitemodels you found are part of MediaPipe’s selfie segmentation suite. It offers both lightweight and heavyweight model options (just like Meet’slite/heavyvariants) optimized for either speed on lower-end devices or higher accuracy. It runs via WASM or WebGL, rendering directly toHTMLCanvasElementexactly as you observed in Meet. - TensorFlow.js Person Segmentation Suite: This includes BodyPix and other specialized segmentation models, all tuned for browser environments. They handle real-time video feed processing, generate a person mask, and let you apply blur, background replacement, or other effects right on the canvas.
- PoseNet (with Segmentation Extensions): While PoseNet is primarily for pose estimation, it can be extended to create person masks for background effects. It’s less focused on this use case than BodyPix or MediaPipe, but still a viable option if you’re already working with pose data.
At their core, all these tools follow the same workflow: perform semantic segmentation to isolate the person in the video feed, then use that mask to apply blur only to the non-person pixels on the HTMLCanvasElement. This keeps all processing local, which is why privacy-focused tools like Meet prefer this approach.
内容的提问来源于stack exchange,提问作者loretoparisi

