Teachable Machine导入p5.js后镜像视频与面部点位错位问题求助
问题:Teachable Machine模型导入p5.js后,镜像视频与面部点位不匹配
尝试将Teachable Machine模型导入p5.js草图,实现视频镜像翻转时,面部点位标记和实际面部位置不匹配。
绘制面部点位的代码片段
push(); textSize(8); if (poser) { //did we get a skeleton yet; for (var i = 0; i < poser.length; i++) { let x = poser[i].position.x; let y = poser[i].position.y; ellipse(x, y, 5, 5); text(poser[i].part, x + 4, y); } } pop();
当前效果

完整p5.js代码
const modelURL = 'https://teachablemachine.withgoogle.com/models/kWOrQ14vz/'; // the json file (model topology) has a reference to the bin file (model weights) const checkpointURL = modelURL + "model.json"; // the metatadata json file contains the text labels of your model and additional information const metadataURL = modelURL + "metadata.json"; const flip = true; // whether to flip the webcam let model; let totalClasses; let myCanvas; let classification = "None Yet"; let probability = "100"; let poser; let video; let flippedVideo; // A function that loads the model from the checkpoint async function load() { model = await tmPose.load(checkpointURL, metadataURL); totalClasses = model.getTotalClasses(); console.log("Number of classes, ", totalClasses); } async function setup() { myCanvas = createCanvas(400, 400); // Call the load function, wait until it finishes loading videoCanvas = createCanvas(320, 240) await load(); video = createCapture(VIDEO, videoReady); video.size(320, 240); video.hide(); } function draw() { background(255); if(video) { flippedVideo = ml5.flipImage(video); image(flippedVideo,0,0,video.width,video.height); } fill(255,0,0) textSize(18); text("Result:" + classification, 10, 40); text("Probability:" + probability, 10, 20) ///ALEX insert if statement here testing classification against apppropriate part of array for this time in your video push(); textSize(8); if (poser) { //did we get a skeleton yet; for (var i = 0; i < poser.length; i++) { let x = poser[i].position.x; let y = poser[i].position.y; ellipse(x, y, 5, 5); text(poser[i].part, x + 4, y); } } pop(); } function videoReady() { console.log("Video Ready"); predict(); } async function predict() { // Prediction #1: run input through posenet // predict can take in an image, video or canvas html element const { pose, posenetOutput } = await model.estimatePose( flippedVideo.elt //webcam.canvas, ); // Prediction 2: run input through teachable machine assification model const prediction = await model.predict( posenetOutput, totalClasses ); // console.log(prediction); // Sort prediction array by probability // So the first classname will have the highest probability const sortedPrediction = prediction.sort((a, b) => -a.probability + b.probability); //communicate these values back to draw function with global variables classification = sortedPrediction[0].className; probability = sortedPrediction[0].probability.toFixed(2); if (pose) poser = pose.keypoints; // is there a skeleton predict(); }
解决方案
问题核心:你给模型传入的是已镜像翻转的视频来预测点位,但绘制时直接使用原始x坐标,未对应镜像转换,导致点位与画面反向。
快速修复方案(推荐)
修改绘制点位的代码,将x坐标转换为镜像后的位置:用视频宽度减去原始x值,让点位和翻转后的画面对齐:
push(); textSize(8); if (poser) { for (var i = 0; i < poser.length; i++) { let originalX = poser[i].position.x; // 计算镜像后的x坐标 let x = video.width - originalX; let y = poser[i].position.y; ellipse(x, y, 5, 5); text(poser[i].part, x + 4, y); } } pop();
备选优化方案
如果不需要模型基于翻转后的画面做分类,可以在predict函数中给model.estimatePose传入原始视频video.elt,只在draw中显示翻转后的画面,再用上述镜像x坐标的方法绘制点位。但此方案会改变模型的分类输入,需根据你的需求选择。
额外提示:draw函数中每次调用ml5.flipImage会创建新图像对象,可能影响性能,可考虑缓存翻转后的图像,或仅在视频帧更新时执行翻转操作。
内容的提问来源于stack exchange,提问作者Skylar W
相关产品推荐
相关产品推荐

