如何将语义分割结果与目标图像叠加,替换分割区域为指定logo
问题描述
我有一张输入图像:
目前已通过以下代码实现了语义分割结果与颜色的叠加:
# sample execution (requires torchvision) from PIL import Image from torchvision import transforms input_image = Image.open("samping.JPG") input_image = input_image.convert("RGB") preprocess = transforms.Compose([ transforms.ToTensor(), transforms.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]), ]) input_tensor = preprocess(input_image) input_batch = input_tensor.unsqueeze(0) # create a mini-batch as expected by the model # move the input and model to GPU for speed if available if torch.cuda.is_available(): input_batch = input_batch.to('cuda') model.to('cuda') with torch.no_grad(): output = model(input_batch)['out'][0] output_predictions = output.argmax(0) # create a color pallette, selecting a color for each class palette = torch.tensor([2 ** 25 - 1, 2 ** 15 - 1, 2 ** 21 - 1]) colors = torch.as_tensor([i for i in range(21)])[:, None] * palette colors = (colors % 255).numpy().astype("uint8") # plot the semantic segmentation predictions of 21 classes in each color r = Image.fromarray(output_predictions.byte().cpu().numpy()).resize(input_image.size) r.putpalette(colors)
得到的颜色叠加分割结果如下:
现在我希望将分割出的区域(如盘子)替换为指定的logo图像:
得到类似目标效果:
请问该如何实现?
实现方案
核心思路是利用分割得到的类别掩码,将logo精准覆盖到目标区域,具体步骤如下:
1. 确定目标类别ID
首先明确要替换区域(比如盘子)对应的语义分割类别索引。以VOC数据集为例,盘子对应diningtable类别,ID为7,需根据实际使用的数据集调整该数值。
2. 生成目标区域掩码
从分割结果中提取目标类别的掩码,并调整至与输入图像一致的尺寸:
# 定义目标类别ID(根据数据集调整) target_class_id = 7 # 生成掩码 mask = (output_predictions == target_class_id).byte().cpu().numpy() mask = Image.fromarray(mask).resize(input_image.size)
3. 处理logo图像
加载logo并调整尺寸,确保带透明通道以实现自然融合:
# 加载logo并转为RGBA格式 logo = Image.open("logo.png").convert("RGBA") # 缩放至输入图像尺寸(或后续按区域大小裁剪) logo = logo.resize(input_image.size, Image.Resampling.LANCZOS)
4. 合成最终图像
使用PIL的paste方法,结合掩码将logo覆盖到原图像的目标区域:
# 将原图像转为RGBA以便混合 result = input_image.convert("RGBA") # 粘贴logo,掩码控制仅在目标区域显示 result.paste(logo, mask=mask) # 转回RGB格式保存或展示 result = result.convert("RGB") result.save("final_result.jpg") result.show()
完整整合代码
将上述逻辑整合到原代码中,完整代码如下:
from PIL import Image from torchvision import transforms import torch input_image = Image.open("samping.JPG") input_image = input_image.convert("RGB") preprocess = transforms.Compose([ transforms.ToTensor(), transforms.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]), ]) input_tensor = preprocess(input_image) input_batch = input_tensor.unsqueeze(0) if torch.cuda.is_available(): input_batch = input_batch.to('cuda') model.to('cuda') with torch.no_grad(): output = model(input_batch)['out'][0] output_predictions = output.argmax(0) # --- 新增区域替换逻辑 --- # 1. 定义目标类别ID target_class_id = 7 # 2. 生成目标区域掩码 mask = (output_predictions == target_class_id).byte().cpu().numpy() mask = Image.fromarray(mask).resize(input_image.size) # 3. 加载并处理logo logo = Image.open("logo.png").convert("RGBA") logo = logo.resize(input_image.size, Image.Resampling.LANCZOS) # 4. 合成图像 result = input_image.convert("RGBA") result.paste(logo, mask=mask) result = result.convert("RGB") # 保存或展示结果 result.save("final_result.jpg") result.show()
优化技巧
如果logo尺寸与目标区域差异较大,可先计算区域边界框,再将logo缩放至对应大小粘贴,效果更自然:
import numpy as np # 计算目标区域的边界框 mask_np = np.array(mask) y_indices, x_indices = np.where(mask_np == 1) x_min, x_max = x_indices.min(), x_indices.max() y_min, y_max = y_indices.min(), y_indices.max() # 缩放logo到区域大小 logo_cropped = logo.resize((x_max - x_min, y_max - y_min), Image.Resampling.LANCZOS) # 粘贴到对应位置 result.paste(logo_cropped, (x_min, y_min), mask=logo_cropped.split()[-1])
内容的提问来源于stack exchange,提问作者confuseman
相关产品推荐
相关产品推荐

