使用OpenCV2提取验证码字符效果不佳,寻求正确操作指导
Hey there! I’ve wrestled with plenty of stubborn CAPTCHAs using OpenCV, so let me share a tried-and-true workflow that usually turns bad extraction results around. More often than not, the issue is missing or misconfigured preprocessing steps—let’s break this down properly:
Captchas are designed to mess with basic image processing, so skipping these steps will almost guarantee bad results:
- Grayscale Conversion: Strip away color channels to simplify processing. Use:
gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY) - Adaptive Thresholding: Forget basic global thresholding—captchas almost always have uneven lighting. Adaptive thresholding adjusts to local pixel values, making characters pop:
We usethresh = cv2.adaptiveThreshold( gray, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C, cv2.THRESH_BINARY_INV, 11, 2 )THRESH_BINARY_INVto invert colors (characters white, background black) which makes contour detection easier. - Noise Reduction: Clean up tiny speckles with morphological operations. An "open" operation (erosion followed by dilation) works great:
For heavier noise, add a median blur before thresholding:kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (2, 2)) cleaned = cv2.morphologyEx(thresh, cv2.MORPH_OPEN, kernel)gray = cv2.medianBlur(gray, 3)
Once your image is clean, you need to isolate individual characters:
- Find External Contours: We only care about outer contours (not inner gaps in letters like 'O' or 'A'):
contours, _ = cv2.findContours(cleaned.copy(), cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE) - Filter Invalid Contours: Get rid of noise and background blobs by checking area and aspect ratio (tweak these values to match your captcha):
valid_contours = [] for cnt in contours: area = cv2.contourArea(cnt) x, y, w, h = cv2.boundingRect(cnt) # Adjust thresholds based on your captcha's character size if 80 < area < 1200 and 0.2 < w/h < 2: valid_contours.append(cnt) - Sort Contours Left-to-Right: Captcha characters are ordered, so sort by the x-coordinate of each contour's bounding box:
valid_contours = sorted(valid_contours, key=lambda x: cv2.boundingRect(x)[0]) - Extract & Standardize Characters: Crop each character and resize to a uniform size (great for feeding into OCR or a model):
extracted_chars = [] for cnt in valid_contours: x, y, w, h = cv2.boundingRect(cnt) char_img = cleaned[y:y+h, x:x+w] # Resize to a consistent size (28x28 is standard for digit/char models) char_img = cv2.resize(char_img, (28, 28)) extracted_chars.append(char_img)
If your captcha has lines, skewed characters, or weird distortions:
- Remove Interference Lines: Target horizontal/vertical lines with specialized kernels:
# Remove horizontal lines kernel_h = cv2.getStructuringElement(cv2.MORPH_RECT, (10, 1)) horizontal_lines = cv2.morphologyEx(thresh, cv2.MORPH_OPEN, kernel_h) no_lines = cv2.subtract(thresh, horizontal_lines) - Skew Correction: If characters are tilted, use
minAreaRectto find the rotation angle and straighten them:rect = cv2.minAreaRect(cnt) angle = rect[2] # Adjust angle for proper rotation (depends on contour orientation) if angle > 45: angle = 90 - angle # Rotate the character (h, w) = char_img.shape[:2] center = (w // 2, h // 2) M = cv2.getRotationMatrix2D(center, angle, 1.0) rotated = cv2.warpAffine(char_img, M, (w, h), flags=cv2.INTER_CUBIC, borderMode=cv2.BORDER_REPLICATE)
After every step, use cv2.imshow("Step Name", img) to check what’s happening. If your contours are picking up noise, tweak the area/aspect ratio thresholds. If characters are merging, adjust the kernel size in morphological operations.
内容的提问来源于stack exchange,提问作者NeebletWorm

