古文字OCR分割:分水岭算法执行后如何获取输出结果?
Hey there! Let's work through your watershed algorithm issues for ancient inscription character segmentation step by step. First, we'll fix some critical bugs in your existing code, then address your two core questions.
先修正原代码中的关键问题
Your current code has a couple of issues that are likely causing the lack of output:
- Mismatched image paths: You first load
templegs.pngfor preprocessing, but later loadtemple.jpgfor watershed processing. These should reference the same image file. - Destructive type conversion: Converting
dist_transformtouint8wipes out the floating-point distance values that are essential for accurately detecting sure foreground areas. Remove this line entirely.
Here's the corrected base code:
import cv2 import numpy as np from matplotlib import pyplot as plt # Load image (use the same path for grayscale and color processing) img_path = 'templegs.png' # Replace with your actual image path gray_img = cv2.imread(img_path, 0) color_img = cv2.imread(img_path, 1) # Preprocessing steps kernel = np.ones((1,1), np.uint8) opening = cv2.morphologyEx(gray_img, cv2.MORPH_OPEN, kernel) blur = cv2.GaussianBlur(opening, (1,1), 0) ret3, th4 = cv2.threshold(blur, 0, 255, cv2.THRESH_BINARY + cv2.THRESH_OTSU) # Sure background area sure_bg = cv2.dilate(opening, kernel, iterations=1) # Finding sure foreground area (removed uint8 conversion) dist_transform = cv2.distanceTransform(opening, cv2.DIST_L2, 3) ret, sure_fg = cv2.threshold(dist_transform, 0.7 * dist_transform.max(), 255, 0) sure_fg = np.uint8(sure_fg) unknown = cv2.subtract(sure_bg, sure_fg) # Marker preparation for watershed ret, markers = cv2.connectedComponents(sure_fg) markers = markers + 1 markers[unknown == 255] = 0 markers = markers.astype('int32') # Apply watershed algorithm markers = cv2.watershed(color_img, markers) color_img[markers == -1] = [255, 0, 0] # Mark watershed boundaries in red
1. 如何显示分水岭算法处理后的输出?
You can use either matplotlib (which you already imported) or OpenCV's native window system to display the processed image. Here are both methods:
Method 1: Using Matplotlib (side-by-side comparison)
Add these lines at the end of your code to view the original image and segmentation result together:
plt.figure(figsize=(12, 6)) # Original image (converted from BGR to RGB for correct Matplotlib display) plt.subplot(121) plt.imshow(cv2.cvtColor(color_img, cv2.COLOR_BGR2RGB)) plt.title('Original Inscription') plt.axis('off') # Watershed segmentation result plt.subplot(122) plt.imshow(cv2.cvtColor(color_img, cv2.COLOR_BGR2RGB)) plt.title('Watershed Segmentation Result') plt.axis('off') plt.show()
Method 2: Using OpenCV's native window
Add these lines to display the result in a standalone window:
cv2.imshow('Watershed Segmentation', color_img) cv2.waitKey(0) # Press any key to close the window cv2.destroyAllWindows()
2. 如何获取分割后的字符作为输出?
To extract individual segmented characters for your OCR system, you'll need to identify each unique marker (excluding boundary and background markers), create a mask for each character, and extract the corresponding Region of Interest (ROI) from the original grayscale image.
Add this code after running the watershed algorithm:
# Get all unique marker values, excluding boundary (-1) and background (1) unique_markers = np.unique(markers) char_markers = [marker for marker in unique_markers if marker not in (-1, 1)] # Extract each character ROI segmented_chars = [] for marker in char_markers: # Create a binary mask for the current character mask = np.where(markers == marker, 255, 0).astype(np.uint8) # Find contours to get the character's bounding box contours, _ = cv2.findContours(mask, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE) if contours: x, y, w, h = cv2.boundingRect(contours[0]) # Extract the character from the grayscale image char_roi = gray_img[y:y+h, x:x+w] segmented_chars.append(char_roi) # Optional: Display all segmented characters plt.figure(figsize=(15, 3)) for i, char in enumerate(segmented_chars): plt.subplot(1, len(segmented_chars), i+1) plt.imshow(char, cmap='gray') plt.title(f'Character {i+1}') plt.axis('off') plt.show()
This code stores all segmented characters in the segmented_chars list, which you can directly feed into your OCR model for further processing.
内容的提问来源于stack exchange,提问作者Venu

