如何使用scikit-image定位图像中加粗编号实现题目图像分割?
Got it, let’s walk through a practical, scikit-image-focused approach to solve this problem—here’s how you can locate those bold, tab-indented number markers (4, 5, 6) and split your image into separate questions:
Step 1: Preprocess the Image to Isolate Text
First, we need to simplify the image so text (especially bold numbers) stands out clearly. Convert to grayscale and apply binary thresholding to separate text from the background:
from skimage import io, color, filters, measure, morphology import numpy as np # Load your target image img = io.imread("your_question_image.png") # Convert to grayscale to reduce complexity gray = color.rgb2gray(img) # Apply binary thresholding (invert if your text is dark on a light background) thresh = filters.threshold_otsu(gray) binary = gray < thresh # Optional: Clean up small noise with morphological opening binary_clean = morphology.opening(binary, morphology.square(3))
Step 2: Focus on the Tab-Indented ROI
Since the numbers are tab-indented, they’ll cluster in a narrow vertical strip on the left side of the image. Crop this region of interest (ROI) to cut down on irrelevant noise:
# Define left ROI (adjust the width percentage to match your image's indentation) roi_width = int(img.shape[1] * 0.1) # Use 10% of the image's width as the left ROI roi = binary_clean[:, :roi_width]
Step 3: Identify Bold Number Regions
Bold numbers have thicker strokes, so their connected pixel regions will be larger than regular text. Use scikit-image’s region properties to filter these out:
# Label connected pixel regions in the ROI labeled_roi = measure.label(roi) # Extract key properties for each region (area, centroid, bounding box) regions = measure.regionprops(labeled_roi) # Filter regions to find bold numbers (adjust the area threshold for your font size) min_bold_area = 100 candidate_regions = [] for region in regions: if region.area >= min_bold_area: candidate_regions.append({ "centroid": region.centroid, # Note: scikit-image uses (y, x) coordinate order "bbox": region.bbox # (min_row, min_col, max_row, max_col) }) # Sort candidates by vertical position (since 4 → 5 → 6 appear top to bottom) candidate_regions.sort(key=lambda x: x["centroid"][0])
Step 4: (Optional) Verify Numbers with OCR
To make sure we’re only targeting 4, 5, 6 (not other bold elements), add a quick OCR check using Tesseract:
import pytesseract from skimage import img_as_ubyte valid_numbers = ["4", "5", "6"] target_regions = [] for region in candidate_regions: min_row, min_col, max_row, max_col = region["bbox"] # Extract the region from the grayscale image num_img = gray[min_row:max_row, min_col:max_col] # Convert to 8-bit format for Tesseract compatibility num_img_8bit = img_as_ubyte(num_img) # Configure Tesseract to only recognize digits text = pytesseract.image_to_string( num_img_8bit, config="--psm 10 --oem 3 -c tessedit_char_whitelist=0123456789" ).strip() if text in valid_numbers: target_regions.append(region) print(f"Found valid number {text} at y-coordinate: {region['centroid'][0]}")
Step 5: Extract Coordinates and Split the Image
Use the vertical (y) coordinates of your target regions to split the original image into three separate questions:
# Get the y-centroid of each target number (use the top of the bounding box instead if preferred) split_ys = [int(region["centroid"][0]) for region in target_regions] # Split the image into individual questions question_1 = img[:split_ys[1], :] # From image top to the start of question 5 question_2 = img[split_ys[1]:split_ys[2], :] # Between question 5 and 6 question_3 = img[split_ys[2]:, :] # From question 6 to image bottom # Save the split images io.imsave("question_1.png", question_1) io.imsave("question_2.png", question_2) io.imsave("question_3.png", question_3)
Key Adjustments for Your Image
- If your image has uneven lighting, replace Otsu thresholding with adaptive thresholding (
filters.threshold_local()). - Tweak
min_bold_areaandroi_widthto match your image’s font size and indentation depth. - Skip the OCR step if you’re confident in region size/position filtering—you can rely on the vertical ordering of candidates to target 4, 5, 6.
内容的提问来源于stack exchange,提问作者Kaushik

