You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:基于Python的三类图像文本分割(Otsu算法失效)

Text Segmentation Fixes for Your Three Image Types

Hey there! Let's work through your text binarization challenges. You noted that Otsu's global thresholding works perfectly for your noise-free third image type, but fails on the other two—let's apply targeted techniques to fix each failure case:

Problem 1: Light Gray Text on the Right Fails to Segment

Your first image has text with lower grayscale values on the right, which gets lost with Otsu's global threshold (since it calculates one threshold for the entire image).

Solution: Adaptive Thresholding

Adaptive thresholding calculates thresholds for small local regions, making it ideal for uneven lighting or varying text contrast. Here's how to implement it:

import cv2

# Read grayscale image
gray = cv2.imread("/home/shrouk/Pictures/f1.png", 0)

# Apply adaptive thresholding (adjust block size and C value based on your image)
# BLOCK_SIZE must be odd; C is subtracted from the local mean to adjust threshold sensitivity
thresholded = cv2.adaptiveThreshold(
    gray, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C, cv2.THRESH_BINARY_INV, 11, 2
)

cv2.imshow("Fixed Light Text Segmentation", thresholded)
cv2.waitKey(0)
cv2.destroyAllWindows()

If adaptive thresholding still leaves some noise, you can add a small Gaussian blur first to smooth the image:

blurred = cv2.GaussianBlur(gray, (3, 3), 0)
thresholded = cv2.adaptiveThreshold(blurred, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C, cv2.THRESH_BINARY_INV, 11, 2)

Problem 2: Dark Background Causes Text Expansion (Bloating)

Otsu's thresholding on dark background images can incorrectly include dark background pixels as part of the text, leading to bloated characters.

Solution: Reverse Thresholding + Morphological Erosion

First, invert the image (since text is likely lighter than the dark background), apply Otsu, then use erosion to shrink the bloated text back to its original shape:

import cv2
import numpy as np

# Read grayscale image
gray = cv2.imread("/home/shrouk/Pictures/f2.png", 0)

# Invert the image (light text on dark background becomes dark text on light)
inverted_gray = cv2.bitwise_not(gray)

# Apply Otsu's thresholding on the inverted image
_, thresholded = cv2.threshold(inverted_gray, 0, 255, cv2.THRESH_BINARY + cv2.THRESH_OTSU)

# Define a small kernel for morphological erosion (adjust size based on bloat severity)
kernel = np.ones((2, 2), np.uint8)
eroded = cv2.erode(thresholded, kernel, iterations=1)

cv2.imshow("Fixed Bloated Text", eroded)
cv2.waitKey(0)
cv2.destroyAllWindows()

Alternatively, you can use adaptive thresholding with inverse binary here too—often it avoids the need for post-processing erosion:

thresholded = cv2.adaptiveThreshold(
    gray, 255, cv2.ADAPTIVE_THRESH_MEAN_C, cv2.THRESH_BINARY, 15, 4
)

For Your Noise-Free Third Image

Your existing Otsu code works great here, but you can keep it as-is or add a quick median blur to ensure no tiny noise creeps in:

gray = cv2.imread("/home/shrouk/Pictures/f3.png", 0)
blurred = cv2.medianBlur(gray, 3)
_, thresholded = cv2.threshold(blurred, 0, 255, cv2.THRESH_BINARY + cv2.THRESH_OTSU)
cv2.imshow("Noise-Free Segmentation", thresholded)

Key Notes

  • Adjust block sizes, C values, and kernel sizes based on your specific image dimensions and text thickness.
  • Always test with your actual images—tweaking these parameters is often needed for perfect results.

内容的提问来源于stack exchange,提问作者shrouk mansour

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:04:26