You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中基于OpenCV轮廓分割数字用于CNN识别的技术问询

Great question! Splitting multi-digit images into individual digits is a common preprocessing step for CNN-based digit recognition, and OpenCV's contour detection is a reliable approach for this task. Let me break down the process step by step, share a working code example, and also cover some alternative methods you can use depending on your use case.

Step-by-Step Guide with OpenCV Contour Detection

1. Preprocess the Image

First, we need to clean up the input image to make contour detection easier. Typical steps include:

  • Convert to grayscale: Reduces the image to a single channel, simplifying processing.
  • Apply thresholding (binarization): Turns the image into black and white, with digits as foreground (usually white) and background as black.
  • Remove noise: Use morphological operations like erosion/dilation to get rid of small artifacts that might be detected as false contours.

2. Detect and Filter Contours

Next, we'll find all contours in the preprocessed image, then filter out any that aren't likely to be digits. Key filters include:

  • Minimum/maximum area: Ignore tiny noise blobs or large background regions.
  • Aspect ratio: Digits tend to have a roughly consistent width-to-height ratio (adjust based on your input images).
  • Bounding rectangle: Ensure the contour is a valid shape that could be a digit.

Critical Note: Don't forget to sort the contours left-to-right! If you skip this, your digits might be in the wrong order, leading to incorrect recognition results (e.g., "8367" instead of "7638").

3. Extract and Normalize Individual Digits

Once we have the valid, sorted contours, we'll extract each digit's region from the original image. For CNN compatibility (especially if you're using a model trained on MNIST), we should:

  • Crop the bounding box around each contour.
  • Resize the cropped digit to a standard size (like 28x28 pixels).
  • Add padding if needed to maintain the digit's aspect ratio (so it doesn't get stretched out of shape).

4. Recognize and Reconstruct the Number

Finally, pass each normalized digit image to your trained CNN classifier, collect the predictions, and concatenate them to get the original number.


Working Code Example

Here's a complete Python script that implements the above steps:

import cv2
import numpy as np

def split_digits(image_path):
    # Read the image
    img = cv2.imread(image_path)
    # Convert to grayscale
    gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
    # Apply adaptive thresholding to handle varying lighting
    thresh = cv2.adaptiveThreshold(gray, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C, cv2.THRESH_BINARY_INV, 11, 2)
    # Remove noise with morphological operations
    kernel = np.ones((2,2), np.uint8)
    cleaned = cv2.morphologyEx(thresh, cv2.MORPH_OPEN, kernel, iterations=1)
    
    # Find contours
    contours, _ = cv2.findContours(cleaned.copy(), cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)
    
    # Filter and sort contours
    digit_contours = []
    for cnt in contours:
        x, y, w, h = cv2.boundingRect(cnt)
        # Filter by area and aspect ratio (adjust these values based on your images)
        if 100 < w*h < 10000 and 0.2 < w/h < 1.0:
            digit_contours.append((x, y, w, h))
    
    # Sort contours left-to-right by x-coordinate
    digit_contours.sort(key=lambda x: x[0])
    
    # Extract and normalize digits
    digits = []
    for (x, y, w, h) in digit_contours:
        # Crop the digit region
        digit_crop = cleaned[y:y+h, x:x+w]
        # Resize to 28x28 (MNIST size) with padding to maintain aspect ratio
        digit_size = 28
        padding = int((h - w)/2) if h > w else 0
        padded = cv2.copyMakeBorder(digit_crop, 0, 0, padding, padding, cv2.BORDER_CONSTANT, value=0)
        resized = cv2.resize(padded, (digit_size, digit_size), interpolation=cv2.INTER_AREA)
        # Normalize pixel values (0-1) for CNN input
        normalized = resized / 255.0
        digits.append(normalized)
    
    return digits

# Example usage
digit_images = split_digits("your_number_image.png")

# Assume you have a trained CNN model called `digit_classifier`
predictions = []
for digit in digit_images:
    # Add batch dimension (CNN expects [batch_size, height, width, channels])
    digit_input = np.expand_dims(digit, axis=(0, -1))
    pred = digit_classifier.predict(digit_input, verbose=0)
    predicted_digit = np.argmax(pred)
    predictions.append(str(predicted_digit))

# Reconstruct the original number
original_number = "".join(predictions)
print(f"Recognized number: {original_number}")

Alternative Split Methods

If contour detection doesn't work well for your specific images (e.g., overlapping digits, complex backgrounds), consider these alternatives:

1. Projection-Based Segmentation

This method uses horizontal and vertical pixel projections to find gaps between digits:

  • Compute the vertical projection (sum of pixels along each column).
  • Identify columns with zero (or near-zero) pixel values—these are the gaps between digits.
  • Split the image at these gap columns to get individual digits.
  • Best for images with well-separated digits and uniform backgrounds.

2. Connected Component Analysis

OpenCV has a built-in function cv2.connectedComponentsWithStats() that groups connected pixels into components. This is similar to contour detection but can be more robust for certain cases:

  • It labels each connected region and provides stats like area, bounding box, etc.
  • You can filter components using the same criteria as contour detection (area, aspect ratio) to isolate digits.

3. Deep Learning-Based Detection

For complex scenarios (e.g., skewed digits, overlapping digits, real-world images), use object detection models like:

  • YOLO (You Only Look Once): Train a small YOLO model to detect individual digits in the image.
  • Faster R-CNN: More accurate but slower than YOLO, good for high-precision needs.
  • This approach handles edge cases better than traditional computer vision methods but requires more training data.

内容的提问来源于stack exchange,提问作者haik

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 08:18:16