Python中基于OpenCV轮廓分割数字用于CNN识别的技术问询
Great question! Splitting multi-digit images into individual digits is a common preprocessing step for CNN-based digit recognition, and OpenCV's contour detection is a reliable approach for this task. Let me break down the process step by step, share a working code example, and also cover some alternative methods you can use depending on your use case.
1. Preprocess the Image
First, we need to clean up the input image to make contour detection easier. Typical steps include:
- Convert to grayscale: Reduces the image to a single channel, simplifying processing.
- Apply thresholding (binarization): Turns the image into black and white, with digits as foreground (usually white) and background as black.
- Remove noise: Use morphological operations like erosion/dilation to get rid of small artifacts that might be detected as false contours.
2. Detect and Filter Contours
Next, we'll find all contours in the preprocessed image, then filter out any that aren't likely to be digits. Key filters include:
- Minimum/maximum area: Ignore tiny noise blobs or large background regions.
- Aspect ratio: Digits tend to have a roughly consistent width-to-height ratio (adjust based on your input images).
- Bounding rectangle: Ensure the contour is a valid shape that could be a digit.
Critical Note: Don't forget to sort the contours left-to-right! If you skip this, your digits might be in the wrong order, leading to incorrect recognition results (e.g., "8367" instead of "7638").
3. Extract and Normalize Individual Digits
Once we have the valid, sorted contours, we'll extract each digit's region from the original image. For CNN compatibility (especially if you're using a model trained on MNIST), we should:
- Crop the bounding box around each contour.
- Resize the cropped digit to a standard size (like 28x28 pixels).
- Add padding if needed to maintain the digit's aspect ratio (so it doesn't get stretched out of shape).
4. Recognize and Reconstruct the Number
Finally, pass each normalized digit image to your trained CNN classifier, collect the predictions, and concatenate them to get the original number.
Working Code Example
Here's a complete Python script that implements the above steps:
import cv2 import numpy as np def split_digits(image_path): # Read the image img = cv2.imread(image_path) # Convert to grayscale gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY) # Apply adaptive thresholding to handle varying lighting thresh = cv2.adaptiveThreshold(gray, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C, cv2.THRESH_BINARY_INV, 11, 2) # Remove noise with morphological operations kernel = np.ones((2,2), np.uint8) cleaned = cv2.morphologyEx(thresh, cv2.MORPH_OPEN, kernel, iterations=1) # Find contours contours, _ = cv2.findContours(cleaned.copy(), cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE) # Filter and sort contours digit_contours = [] for cnt in contours: x, y, w, h = cv2.boundingRect(cnt) # Filter by area and aspect ratio (adjust these values based on your images) if 100 < w*h < 10000 and 0.2 < w/h < 1.0: digit_contours.append((x, y, w, h)) # Sort contours left-to-right by x-coordinate digit_contours.sort(key=lambda x: x[0]) # Extract and normalize digits digits = [] for (x, y, w, h) in digit_contours: # Crop the digit region digit_crop = cleaned[y:y+h, x:x+w] # Resize to 28x28 (MNIST size) with padding to maintain aspect ratio digit_size = 28 padding = int((h - w)/2) if h > w else 0 padded = cv2.copyMakeBorder(digit_crop, 0, 0, padding, padding, cv2.BORDER_CONSTANT, value=0) resized = cv2.resize(padded, (digit_size, digit_size), interpolation=cv2.INTER_AREA) # Normalize pixel values (0-1) for CNN input normalized = resized / 255.0 digits.append(normalized) return digits # Example usage digit_images = split_digits("your_number_image.png") # Assume you have a trained CNN model called `digit_classifier` predictions = [] for digit in digit_images: # Add batch dimension (CNN expects [batch_size, height, width, channels]) digit_input = np.expand_dims(digit, axis=(0, -1)) pred = digit_classifier.predict(digit_input, verbose=0) predicted_digit = np.argmax(pred) predictions.append(str(predicted_digit)) # Reconstruct the original number original_number = "".join(predictions) print(f"Recognized number: {original_number}")
If contour detection doesn't work well for your specific images (e.g., overlapping digits, complex backgrounds), consider these alternatives:
1. Projection-Based Segmentation
This method uses horizontal and vertical pixel projections to find gaps between digits:
- Compute the vertical projection (sum of pixels along each column).
- Identify columns with zero (or near-zero) pixel values—these are the gaps between digits.
- Split the image at these gap columns to get individual digits.
- Best for images with well-separated digits and uniform backgrounds.
2. Connected Component Analysis
OpenCV has a built-in function cv2.connectedComponentsWithStats() that groups connected pixels into components. This is similar to contour detection but can be more robust for certain cases:
- It labels each connected region and provides stats like area, bounding box, etc.
- You can filter components using the same criteria as contour detection (area, aspect ratio) to isolate digits.
3. Deep Learning-Based Detection
For complex scenarios (e.g., skewed digits, overlapping digits, real-world images), use object detection models like:
- YOLO (You Only Look Once): Train a small YOLO model to detect individual digits in the image.
- Faster R-CNN: More accurate but slower than YOLO, good for high-precision needs.
- This approach handles edge cases better than traditional computer vision methods but requires more training data.
内容的提问来源于stack exchange,提问作者haik

