You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在文本行间添加白线解决Tesseract OCR的跨行字符粘连问题?

Fixing Cross-Line Character Adhesion in Tesseract OCR for Long Receipts

Absolutely, adding white lines or blocks between text lines is a smart pre-processing fix for this exact issue—here’s how to implement it effectively, along with complementary tweaks for Tesseract:

Why This Works

Tesseract’s line segmentation relies on clear visual separation between lines. Those stuck characters (like your p and M) confuse the engine into merging lines, so inserting a thin white barrier breaks those connections and helps Tesseract correctly identify individual text rows.

Step 1: Pre-Process the Image to Insert White Lines

If you’re using Python (the most common setup for Tesseract workflows), tools like OpenCV make this straightforward. You can either use fixed intervals (great for consistent line spacing, which your high-quality scan likely has) or detect text lines dynamically for uneven spacing.

Example with Fixed Intervals (Python + OpenCV)

import cv2
import numpy as np

# Load your scanned receipt image
img = cv2.imread("long_receipt.jpg")
gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
height, width = gray.shape

# Adjust these values based on your receipt's text size:
line_height = 40  # Approximate height of one text line
separator_height = 3  # Thin white line to break粘连

modified_segments = []
for y in range(0, height, line_height):
    # Grab the current text line segment
    line_segment = gray[y:y+line_height, :]
    modified_segments.append(line_segment)
    
    # Add a white line if we're not at the last line
    if y + line_height < height:
        white_line = np.ones((separator_height, width), dtype=np.uint8) * 255
        modified_segments.append(white_line)

# Combine all segments into the final image
final_img = np.vstack(modified_segments)
cv2.imwrite("receipt_with_separators.jpg", final_img)

After saving the modified image, run Tesseract as usual:

tesseract receipt_with_separators.jpg ocr_output

For Uneven Line Spacing

If your receipt has inconsistent line gaps, use contour detection to find the bounding boxes of each text line first. Calculate the space between consecutive boxes, then insert white lines in those gaps. This takes a bit more code but adapts to variable spacing.

Step 2: Tesseract Configuration Tweaks

Pair the image pre-processing with these Tesseract settings to boost accuracy further:

  • --psm 6: Tells Tesseract to treat the input as a single uniform block of text (perfect for receipts).
  • -c textord_min_xheight=20: Adjust this value to match your text’s minimum height—helps Tesseract avoid merging small characters across lines.
  • -c textord_noise_rejection=1: Reduces minor noise that might interfere with line segmentation.

Pro Tips

  • Start with a thin separator (2-3 pixels) first—too thick can make Tesseract misclassify the white space as a separate block.
  • Test with a small section of your receipt first to tweak line_height and separator_height to your specific scan.

内容的提问来源于stack exchange,提问作者Ahmed Bilal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:10:52