You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将语义分割结果与目标图像叠加,替换分割区域为指定logo

问题描述

我有一张输入图像:
未分割的图像

目前已通过以下代码实现了语义分割结果与颜色的叠加:

# sample execution (requires torchvision)
from PIL import Image
from torchvision import transforms
input_image = Image.open("samping.JPG")
input_image = input_image.convert("RGB")
preprocess = transforms.Compose([
    transforms.ToTensor(),
    transforms.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),
])

input_tensor = preprocess(input_image)
input_batch = input_tensor.unsqueeze(0) # create a mini-batch as expected by the model

# move the input and model to GPU for speed if available
if torch.cuda.is_available():
    input_batch = input_batch.to('cuda')
    model.to('cuda')

with torch.no_grad():
    output = model(input_batch)['out'][0]
output_predictions = output.argmax(0)

# create a color pallette, selecting a color for each class
palette = torch.tensor([2 ** 25 - 1, 2 ** 15 - 1, 2 ** 21 - 1])
colors = torch.as_tensor([i for i in range(21)])[:, None] * palette
colors = (colors % 255).numpy().astype("uint8")

# plot the semantic segmentation predictions of 21 classes in each color
r = Image.fromarray(output_predictions.byte().cpu().numpy()).resize(input_image.size)
r.putpalette(colors)

得到的颜色叠加分割结果如下:
带颜色的分割图像

现在我希望将分割出的区域(如盘子)替换为指定的logo图像:
用于叠加的图像

得到类似目标效果:
期望的分割图像

请问该如何实现?


实现方案

核心思路是利用分割得到的类别掩码,将logo精准覆盖到目标区域,具体步骤如下:

1. 确定目标类别ID

首先明确要替换区域(比如盘子)对应的语义分割类别索引。以VOC数据集为例,盘子对应diningtable类别,ID为7,需根据实际使用的数据集调整该数值。

2. 生成目标区域掩码

从分割结果中提取目标类别的掩码,并调整至与输入图像一致的尺寸:

# 定义目标类别ID(根据数据集调整)
target_class_id = 7
# 生成掩码
mask = (output_predictions == target_class_id).byte().cpu().numpy()
mask = Image.fromarray(mask).resize(input_image.size)

3. 处理logo图像

加载logo并调整尺寸,确保带透明通道以实现自然融合:

# 加载logo并转为RGBA格式
logo = Image.open("logo.png").convert("RGBA")
# 缩放至输入图像尺寸(或后续按区域大小裁剪)
logo = logo.resize(input_image.size, Image.Resampling.LANCZOS)

4. 合成最终图像

使用PIL的paste方法,结合掩码将logo覆盖到原图像的目标区域:

# 将原图像转为RGBA以便混合
result = input_image.convert("RGBA")
# 粘贴logo,掩码控制仅在目标区域显示
result.paste(logo, mask=mask)
# 转回RGB格式保存或展示
result = result.convert("RGB")
result.save("final_result.jpg")
result.show()

完整整合代码

将上述逻辑整合到原代码中,完整代码如下:

from PIL import Image
from torchvision import transforms
import torch

input_image = Image.open("samping.JPG")
input_image = input_image.convert("RGB")
preprocess = transforms.Compose([
    transforms.ToTensor(),
    transforms.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),
])

input_tensor = preprocess(input_image)
input_batch = input_tensor.unsqueeze(0)

if torch.cuda.is_available():
    input_batch = input_batch.to('cuda')
    model.to('cuda')

with torch.no_grad():
    output = model(input_batch)['out'][0]
output_predictions = output.argmax(0)

# --- 新增区域替换逻辑 ---
# 1. 定义目标类别ID
target_class_id = 7
# 2. 生成目标区域掩码
mask = (output_predictions == target_class_id).byte().cpu().numpy()
mask = Image.fromarray(mask).resize(input_image.size)
# 3. 加载并处理logo
logo = Image.open("logo.png").convert("RGBA")
logo = logo.resize(input_image.size, Image.Resampling.LANCZOS)
# 4. 合成图像
result = input_image.convert("RGBA")
result.paste(logo, mask=mask)
result = result.convert("RGB")
# 保存或展示结果
result.save("final_result.jpg")
result.show()

优化技巧

如果logo尺寸与目标区域差异较大,可先计算区域边界框,再将logo缩放至对应大小粘贴,效果更自然:

import numpy as np
# 计算目标区域的边界框
mask_np = np.array(mask)
y_indices, x_indices = np.where(mask_np == 1)
x_min, x_max = x_indices.min(), x_indices.max()
y_min, y_max = y_indices.min(), y_indices.max()
# 缩放logo到区域大小
logo_cropped = logo.resize((x_max - x_min, y_max - y_min), Image.Resampling.LANCZOS)
# 粘贴到对应位置
result.paste(logo_cropped, (x_min, y_min), mask=logo_cropped.split()[-1])

内容的提问来源于stack exchange,提问作者confuseman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 15:27:57