You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

双背景(黑白)图像全文本提取方案咨询:现有OCR代码仅识别白色背景文本的优化需求

Solution for Extracting Text from Images with Mixed Black/White Backgrounds

Your current code only targets one text-background pairing (white background + dark text), which is why you're missing the text that sits on black backgrounds. The fix here is to handle both contrast scenarios separately, then combine the results to capture all text. Let's walk through the adjusted approach:

Step-by-Step Breakdown

  1. Switch to Grayscale: Grayscale simplifies working with black/white contrast, avoiding the complexity of HSV color space for this specific problem.
  2. Dual Thresholding:
    • First, isolate dark text on white backgrounds using a binary inverse threshold.
    • Second, isolate light text on black backgrounds using a standard binary threshold (after implicitly inverting the contrast).
  3. Combine Processed Images: Merge the two thresholded outputs to create a single image where all text is visible against a uniform background, then run OCR on this combined result.

Modified Code

import cv2
import pytesseract
import numpy as np

# Tesseract configuration (adjust paths/config to match your setup)
# pytesseract.pytesseract.tesseract_cmd = r"C:\Program Files\Tesseract-OCR\tesseract.exe"
tessdata_dir_config = '--tessdata-dir "/usr/local/Cellar/tesseract-lang/4.1.0/share/tessdata"'

# Load image and convert to grayscale
image = cv2.imread('/Users/snrt1/PycharmProjects/pythonProjectopencv/image/150.jpg')
gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY)

# 1. Extract dark text on white backgrounds
_, thresh_white_bg = cv2.threshold(gray, 127, 255, cv2.THRESH_BINARY_INV)

# 2. Extract light text on black backgrounds
_, thresh_black_bg = cv2.threshold(gray, 127, 255, cv2.THRESH_BINARY)

# Combine both results to get all text in one image
combined_text = cv2.bitwise_or(thresh_white_bg, thresh_black_bg)

# Optional: Reduce noise to improve OCR accuracy
combined_text = cv2.GaussianBlur(combined_text, (3, 3), 0)

# Run OCR with your specified language and layout config
full_text = pytesseract.image_to_string(combined_text, lang='eng+ara', config='--psm 6')
print(full_text)

# Debug windows to visualize each step
cv2.imshow('Original Grayscale', gray)
cv2.imshow('Text on White BG', thresh_white_bg)
cv2.imshow('Text on Black BG', thresh_black_bg)
cv2.imshow('Combined Text Image', combined_text)
cv2.waitKey(0)
cv2.destroyAllWindows()

Quick Tips for Better Results

  • Tweak Threshold Value: If 127 doesn't capture text well, experiment with values between 100-150 to match your image's lighting.
  • Adjust PSM Mode: --psm 6 works for uniform text blocks. For images with scattered text, try --psm 3 (default) or --psm 11.
  • Noise Removal: Adding a median blur (cv2.medianBlur(combined_text, 3)) instead of Gaussian can help with salt-and-pepper noise.

内容的提问来源于stack exchange,提问作者Khalid Elfayq

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 06:47:38