PyTorch中DSB问题预测结果转Run Length Encoding格式求助
Got it, let's tackle this RLE conversion for your DSB submission—this is such a common step once you've got your segmentation masks ready from PyTorch. I'll break down the process clearly, with code examples you can copy-paste and adapt.
First, Let's Recap the DSB RLE Rules
The event uses RLE to reduce submission file size: instead of submitting all pixel indices for segmentation results, you only need pairs of values where the first number is the 1-based starting pixel position and the second is the length of consecutive foreground pixels. For example,
1 3means 3 consecutive foreground pixels starting at the 1st pixel (row-major order: left to right, top to bottom).
Step 1: Prepare Your PyTorch Mask
First, you'll need to convert your PyTorch tensor output into a usable 2D binary mask:
- If your model outputs a probability map (e.g., from
sigmoid()), threshold it to get 0 (background) and 1 (foreground) values. - Convert the tensor to a numpy array, and adjust dimensions (e.g., from
[1, H, W]to[H, W]by squeezing the channel dimension).
Example code for this step:
import torch import numpy as np # Assume pred is your PyTorch model output (shape: [1, H, W] for single channel) pred_tensor = torch.load("your_prediction.pt") # Or get directly from model forward pass # Threshold to get binary mask binary_mask = (pred_tensor.squeeze().cpu().numpy() > 0.5).astype(np.int32)
Step 2: Single Instance RLE Conversion
If your mask has only one foreground object, use this function to convert it to RLE:
def mask_to_rle(binary_mask): # binary_mask: 2D numpy array (H x W), values 0 or 1 # Flatten mask in row-major order (matches DSB's pixel counting) flat_mask = binary_mask.flatten(order='C') # Pad mask with 0s at start and end to catch edge cases (e.g., mask starts at pixel 1) padded_mask = np.concatenate([[0], flat_mask, [0]]) # Find indices where the mask value changes (0→1 or 1→0) change_indices = np.where(padded_mask[1:] != padded_mask[:-1])[0] # Generate RLE pairs: (start position, length) rle_pairs = [] for i in range(0, len(change_indices), 2): start_pos = change_indices[i] + 1 # Convert to 1-based index length = change_indices[i+1] - change_indices[i] rle_pairs.extend([str(start_pos), str(length)]) # Return empty string as '0' if no foreground pixels (per DSB requirements) return ' '.join(rle_pairs) if rle_pairs else '0'
Step 3: Multi-Instance RLE Conversion
If your mask has multiple separate foreground objects, first label each connected component, then generate RLE for each instance:
from skimage.measure import label def multi_instance_mask_to_rle(binary_mask): # Label each unique connected foreground component labeled_mask = label(binary_mask, connectivity=2) # 2=8-connectivity, adjust if needed rle_list = [] # Iterate over each unique object (skip background label 0) for obj_label in np.unique(labeled_mask): if obj_label == 0: continue # Extract mask for this specific object obj_mask = (labeled_mask == obj_label).astype(np.int32) # Convert to RLE rle_list.append(mask_to_rle(obj_mask)) return rle_list
Step 4: Generate Submission File
DSB's submission format is a CSV with two columns: ImageId and EncodedPixels. Each foreground instance gets its own row (even for the same image). Example code to build the CSV:
import pandas as pd # Assume you have a list of image IDs and their corresponding masks submission_data = [] for image_id, mask in zip(your_image_ids, your_masks_list): rles = multi_instance_mask_to_rle(mask) if not rles: # No foreground, add a single row with '0' submission_data.append({"ImageId": image_id, "EncodedPixels": "0"}) else: for rle in rles: submission_data.append({"ImageId": image_id, "EncodedPixels": rle}) # Save to CSV submission_df = pd.DataFrame(submission_data) submission_df.to_csv("dsb_submission.csv", index=False)
Key Notes to Avoid Mistakes
- 1-based vs 0-based indices: DSB uses 1-based pixel numbering—don't forget to add 1 to your start positions!
- Row-major order: Make sure you flatten the mask in row-major (left→right, top→bottom) order, which is the default for
flatten(order='C'). - Empty masks: Always return
'0'for images with no foreground, instead of an empty string.
Hope this helps you get your submission sorted smoothly! If you run into edge cases (like tiny objects or weird mask shapes), feel free to tweak the code to fit your specific predictions.
内容的提问来源于stack exchange,提问作者Ryan

