You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyTorch中DSB问题预测结果转Run Length Encoding格式求助

Converting PyTorch Segmentation Predictions to DSB's Run Length Encoding (RLE)

Got it, let's tackle this RLE conversion for your DSB submission—this is such a common step once you've got your segmentation masks ready from PyTorch. I'll break down the process clearly, with code examples you can copy-paste and adapt.

First, Let's Recap the DSB RLE Rules

The event uses RLE to reduce submission file size: instead of submitting all pixel indices for segmentation results, you only need pairs of values where the first number is the 1-based starting pixel position and the second is the length of consecutive foreground pixels. For example, 1 3 means 3 consecutive foreground pixels starting at the 1st pixel (row-major order: left to right, top to bottom).

Step 1: Prepare Your PyTorch Mask

First, you'll need to convert your PyTorch tensor output into a usable 2D binary mask:

  • If your model outputs a probability map (e.g., from sigmoid()), threshold it to get 0 (background) and 1 (foreground) values.
  • Convert the tensor to a numpy array, and adjust dimensions (e.g., from [1, H, W] to [H, W] by squeezing the channel dimension).

Example code for this step:

import torch
import numpy as np

# Assume pred is your PyTorch model output (shape: [1, H, W] for single channel)
pred_tensor = torch.load("your_prediction.pt")  # Or get directly from model forward pass
# Threshold to get binary mask
binary_mask = (pred_tensor.squeeze().cpu().numpy() > 0.5).astype(np.int32)

Step 2: Single Instance RLE Conversion

If your mask has only one foreground object, use this function to convert it to RLE:

def mask_to_rle(binary_mask):
    # binary_mask: 2D numpy array (H x W), values 0 or 1
    # Flatten mask in row-major order (matches DSB's pixel counting)
    flat_mask = binary_mask.flatten(order='C')
    # Pad mask with 0s at start and end to catch edge cases (e.g., mask starts at pixel 1)
    padded_mask = np.concatenate([[0], flat_mask, [0]])
    # Find indices where the mask value changes (0→1 or 1→0)
    change_indices = np.where(padded_mask[1:] != padded_mask[:-1])[0]
    # Generate RLE pairs: (start position, length)
    rle_pairs = []
    for i in range(0, len(change_indices), 2):
        start_pos = change_indices[i] + 1  # Convert to 1-based index
        length = change_indices[i+1] - change_indices[i]
        rle_pairs.extend([str(start_pos), str(length)])
    # Return empty string as '0' if no foreground pixels (per DSB requirements)
    return ' '.join(rle_pairs) if rle_pairs else '0'

Step 3: Multi-Instance RLE Conversion

If your mask has multiple separate foreground objects, first label each connected component, then generate RLE for each instance:

from skimage.measure import label

def multi_instance_mask_to_rle(binary_mask):
    # Label each unique connected foreground component
    labeled_mask = label(binary_mask, connectivity=2)  # 2=8-connectivity, adjust if needed
    rle_list = []
    # Iterate over each unique object (skip background label 0)
    for obj_label in np.unique(labeled_mask):
        if obj_label == 0:
            continue
        # Extract mask for this specific object
        obj_mask = (labeled_mask == obj_label).astype(np.int32)
        # Convert to RLE
        rle_list.append(mask_to_rle(obj_mask))
    return rle_list

Step 4: Generate Submission File

DSB's submission format is a CSV with two columns: ImageId and EncodedPixels. Each foreground instance gets its own row (even for the same image). Example code to build the CSV:

import pandas as pd

# Assume you have a list of image IDs and their corresponding masks
submission_data = []
for image_id, mask in zip(your_image_ids, your_masks_list):
    rles = multi_instance_mask_to_rle(mask)
    if not rles:
        # No foreground, add a single row with '0'
        submission_data.append({"ImageId": image_id, "EncodedPixels": "0"})
    else:
        for rle in rles:
            submission_data.append({"ImageId": image_id, "EncodedPixels": rle})

# Save to CSV
submission_df = pd.DataFrame(submission_data)
submission_df.to_csv("dsb_submission.csv", index=False)

Key Notes to Avoid Mistakes

  • 1-based vs 0-based indices: DSB uses 1-based pixel numbering—don't forget to add 1 to your start positions!
  • Row-major order: Make sure you flatten the mask in row-major (left→right, top→bottom) order, which is the default for flatten(order='C').
  • Empty masks: Always return '0' for images with no foreground, instead of an empty string.

Hope this helps you get your submission sorted smoothly! If you run into edge cases (like tiny objects or weird mask shapes), feel free to tweak the code to fit your specific predictions.

内容的提问来源于stack exchange,提问作者Ryan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 03:30:44