You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

古文字OCR分割:分水岭算法执行后如何获取输出结果?

解决分水岭算法的图像显示与古铭文字符分割输出问题

Hey there! Let's work through your watershed algorithm issues for ancient inscription character segmentation step by step. First, we'll fix some critical bugs in your existing code, then address your two core questions.

先修正原代码中的关键问题

Your current code has a couple of issues that are likely causing the lack of output:

  • Mismatched image paths: You first load templegs.png for preprocessing, but later load temple.jpg for watershed processing. These should reference the same image file.
  • Destructive type conversion: Converting dist_transform to uint8 wipes out the floating-point distance values that are essential for accurately detecting sure foreground areas. Remove this line entirely.

Here's the corrected base code:

import cv2
import numpy as np
from matplotlib import pyplot as plt

# Load image (use the same path for grayscale and color processing)
img_path = 'templegs.png'  # Replace with your actual image path
gray_img = cv2.imread(img_path, 0)
color_img = cv2.imread(img_path, 1)

# Preprocessing steps
kernel = np.ones((1,1), np.uint8)
opening = cv2.morphologyEx(gray_img, cv2.MORPH_OPEN, kernel)
blur = cv2.GaussianBlur(opening, (1,1), 0)
ret3, th4 = cv2.threshold(blur, 0, 255, cv2.THRESH_BINARY + cv2.THRESH_OTSU)

# Sure background area
sure_bg = cv2.dilate(opening, kernel, iterations=1)

# Finding sure foreground area (removed uint8 conversion)
dist_transform = cv2.distanceTransform(opening, cv2.DIST_L2, 3)
ret, sure_fg = cv2.threshold(dist_transform, 0.7 * dist_transform.max(), 255, 0)
sure_fg = np.uint8(sure_fg)

unknown = cv2.subtract(sure_bg, sure_fg)

# Marker preparation for watershed
ret, markers = cv2.connectedComponents(sure_fg)
markers = markers + 1
markers[unknown == 255] = 0
markers = markers.astype('int32')

# Apply watershed algorithm
markers = cv2.watershed(color_img, markers)
color_img[markers == -1] = [255, 0, 0]  # Mark watershed boundaries in red

1. 如何显示分水岭算法处理后的输出?

You can use either matplotlib (which you already imported) or OpenCV's native window system to display the processed image. Here are both methods:

Method 1: Using Matplotlib (side-by-side comparison)

Add these lines at the end of your code to view the original image and segmentation result together:

plt.figure(figsize=(12, 6))

# Original image (converted from BGR to RGB for correct Matplotlib display)
plt.subplot(121)
plt.imshow(cv2.cvtColor(color_img, cv2.COLOR_BGR2RGB))
plt.title('Original Inscription')
plt.axis('off')

# Watershed segmentation result
plt.subplot(122)
plt.imshow(cv2.cvtColor(color_img, cv2.COLOR_BGR2RGB))
plt.title('Watershed Segmentation Result')
plt.axis('off')

plt.show()

Method 2: Using OpenCV's native window

Add these lines to display the result in a standalone window:

cv2.imshow('Watershed Segmentation', color_img)
cv2.waitKey(0)  # Press any key to close the window
cv2.destroyAllWindows()

2. 如何获取分割后的字符作为输出?

To extract individual segmented characters for your OCR system, you'll need to identify each unique marker (excluding boundary and background markers), create a mask for each character, and extract the corresponding Region of Interest (ROI) from the original grayscale image.

Add this code after running the watershed algorithm:

# Get all unique marker values, excluding boundary (-1) and background (1)
unique_markers = np.unique(markers)
char_markers = [marker for marker in unique_markers if marker not in (-1, 1)]

# Extract each character ROI
segmented_chars = []
for marker in char_markers:
    # Create a binary mask for the current character
    mask = np.where(markers == marker, 255, 0).astype(np.uint8)
    # Find contours to get the character's bounding box
    contours, _ = cv2.findContours(mask, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)
    if contours:
        x, y, w, h = cv2.boundingRect(contours[0])
        # Extract the character from the grayscale image
        char_roi = gray_img[y:y+h, x:x+w]
        segmented_chars.append(char_roi)

# Optional: Display all segmented characters
plt.figure(figsize=(15, 3))
for i, char in enumerate(segmented_chars):
    plt.subplot(1, len(segmented_chars), i+1)
    plt.imshow(char, cmap='gray')
    plt.title(f'Character {i+1}')
    plt.axis('off')
plt.show()

This code stores all segmented characters in the segmented_chars list, which you can directly feed into your OCR model for further processing.

内容的提问来源于stack exchange,提问作者Venu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 05:37:51