You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

寻求可生成带标签大图像分块(Patches)及对应标签的Python/PyTorch库:有丝分裂目标检测场景技术求助

Hey there, I’ve dealt with exactly this problem for large medical image datasets (including mitosis detection!), so let’s break down how to solve this properly. The core challenge here is ensuring that every spatial operation you apply to your large image gets mirrored exactly on your labels—whether those are bounding boxes or segmentation masks. Below are practical, PyTorch-compatible solutions:

Solution for Synchronous Image and Label Patching

1. Manual Implementation (Full Control)

If you want custom logic tailored to your dataset, this approach gives you complete oversight. We’ll cover both bounding box and segmentation mask labels:

For Bounding Box Labels

Assuming your labels are in [x1, y1, x2, y2] format (Pascal VOC style), here’s how to crop patches and adjust bboxes to local patch coordinates:

import numpy as np
from PIL import Image

def split_image_and_bboxes(image, bboxes, patch_size=(512,512), stride=512):
    patches = []
    patch_bboxes = []
    img_h, img_w = image.shape[:2]
    
    # Slide window across the image
    for y in range(0, img_h - patch_size[0] + 1, stride):
        for x in range(0, img_w - patch_size[1] + 1, stride):
            # Extract image patch
            patch = image[y:y+patch_size[0], x:x+patch_size[1]]
            patches.append(patch)
            
            # Adjust bboxes to patch-local coordinates
            adjusted_bboxes = []
            for bbox in bboxes:
                x1, y1, x2, y2 = bbox
                # Check if the bbox overlaps with the current patch
                overlap_x1 = max(x1, x)
                overlap_y1 = max(y1, y)
                overlap_x2 = min(x2, x + patch_size[1])
                overlap_y2 = min(y2, y + patch_size[0])
                
                if overlap_x2 > overlap_x1 and overlap_y2 > overlap_y1:
                    # Convert to patch-specific coordinates
                    adj_x1 = overlap_x1 - x
                    adj_y1 = overlap_y1 - y
                    adj_x2 = overlap_x2 - x
                    adj_y2 = overlap_y2 - y
                    adjusted_bboxes.append([adj_x1, adj_y1, adj_x2, adj_y2])
            
            patch_bboxes.append(adjusted_bboxes)
    
    return patches, patch_bboxes

# Example usage
image = np.array(Image.open("large_mitosis_image.jpg"))
bboxes = [[100,200,150,250], [300,400,350,450]]  # Load your actual bboxes here
patches, patch_bboxes = split_image_and_bboxes(image, bboxes, patch_size=(512,512), stride=512)

For Segmentation Mask Labels

If your labels are pixel-wise masks (same dimensions as the image), cropping is straightforward—just apply the same slice to both image and mask:

def split_image_and_mask(image, mask, patch_size=(512,512), stride=512):
    patches = []
    mask_patches = []
    img_h, img_w = image.shape[:2]
    
    for y in range(0, img_h - patch_size[0] + 1, stride):
        for x in range(0, img_w - patch_size[1] + 1, stride):
            patch = image[y:y+patch_size[0], x:x+patch_size[1]]
            mask_patch = mask[y:y+patch_size[0], x:x+patch_size[1]]
            patches.append(patch)
            mask_patches.append(mask_patch)
    
    return patches, mask_patches

2. Using Albumentations (PyTorch-Friendly Library)

Albumentations is the gold standard for synchronized image/label transformations in computer vision. It automatically handles bounding boxes, masks, and keypoints with minimal code:

import albumentations as A
from PIL import Image
import numpy as np

# Define your crop transform (use RandomCrop for random patches or Crop for fixed sliding window)
transform = A.Compose([
    A.RandomCrop(width=512, height=512),  # Replace with A.Crop for fixed coordinates
], bbox_params=A.BboxParams(format='pascal_voc', label_fields=['labels']))

# Load your data
image = np.array(Image.open("large_mitosis_image.jpg"))
bboxes = [[100,200,150,250], [300,400,350,450]]
labels = [0, 0]  # Class ID for mitosis

# Generate multiple patches
num_patches = 10
patches = []
patch_bboxes = []
patch_labels = []

for _ in range(num_patches):
    augmented = transform(image=image, bboxes=bboxes, labels=labels)
    patches.append(augmented['image'])
    patch_bboxes.append(augmented['bboxes'])
    patch_labels.append(augmented['labels'])

For sliding window patches, loop through fixed x/y coordinates and use A.Crop(top=y, left=x, height=512, width=512) instead of RandomCrop.

3. TorchVision Transforms (With Custom Wrappers)

TorchVision works well for images, but you’ll need to wrap transforms to handle labels. Here’s how to use torchvision.transforms.functional for cropping:

import torchvision.transforms.functional as F
from PIL import Image

def crop_image_and_bboxes(image, bboxes, top, left, height, width):
    cropped_img = F.crop(image, top, left, height, width)
    adjusted_bboxes = []
    for x1, y1, x2, y2 in bboxes:
        # Adjust bbox coordinates to patch space
        adj_x1 = max(0, x1 - left)
        adj_y1 = max(0, y1 - top)
        adj_x2 = min(width, x2 - left)
        adj_y2 = min(height, y2 - top)
        if adj_x2 > adj_x1 and adj_y2 > adj_y1:
            adjusted_bboxes.append([adj_x1, adj_y1, adj_x2, adj_y2])
    return cropped_img, adjusted_bboxes

Key Tips

  • Overlap Handling: If using a stride smaller than patch size (for overlapping patches), the manual code above already handles bbox overlap correctly.
  • Edge Patches: For patches at the bottom/right edge that don’t fit the full patch size, use padding (e.g., A.PadIfNeeded in Albumentations) to avoid losing data.
  • Label Format: Always double-check that your bbox format matches what the library expects (Pascal VOC, COCO, etc.).

内容的提问来源于stack exchange,提问作者Ixion Chowdhury

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 06:51:41