You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何定位图像中新增的普通/弯曲/旋转文本区域并绘制包围框

Hey there! The problem with your current code is that it uses basic contour detection on pixel differences, which falls flat for rotated or curved text—those often have scattered diff pixels, and axis-aligned rectangles don’t fit their actual shape. Let’s solve this by pairing text detection models (built for arbitrary text shapes) with your existing image diff logic to zero in on exactly the new text regions.

Solution Overview

We’ll use a pre-trained text detection model (CRAFT, which excels at curved/rotated text) alongside image difference filtering. This ensures we only target newly added text, not existing content or non-text differences, and draw precise bounding polygons instead of rigid rectangles.

Step 1: Install Dependencies

First, grab the required packages:

pip install opencv-python numpy scikit-image craft-text-detector

Step 2: Full Implementation Code

This code combines diff analysis with CRAFT to detect and draw accurate bounding shapes for all text types:

import cv2
import numpy as np
from skimage.metrics import structural_similarity as compare_ssim
from craft_text_detector import Craft

def detect_new_text(before_img_path, after_img_path):
    # Load and preprocess images
    before = cv2.imread(before_img_path)
    after = cv2.imread(after_img_path)
    before_gray = cv2.cvtColor(before, cv2.COLOR_BGR2GRAY)
    after_gray = cv2.cvtColor(after, cv2.COLOR_BGR2GRAY)

    # Generate difference mask using SSIM
    (_, diff) = compare_ssim(before_gray, after_gray, full=True)
    diff = (diff * 255).astype("uint8")
    thresh = cv2.threshold(diff, 0, 255, cv2.THRESH_BINARY_INV | cv2.THRESH_OTSU)[1]

    # Initialize CRAFT text detector (handles curved/rotated text)
    craft = Craft(output_dir=None, crop_type="poly", cuda=False)
    text_preds = craft.detect_text(after)
    all_text_polygons = text_preds["polygons"]

    # Filter polygons to only those in the difference region (new text)
    new_text_polygons = []
    for poly in all_text_polygons:
        # Create mask for current text polygon
        poly_mask = np.zeros_like(thresh)
        cv2.fillPoly(poly_mask, [np.array(poly, dtype=np.int32)], 255)
        # Check if polygon overlaps significantly with diff mask
        overlap = cv2.bitwise_and(thresh, poly_mask)
        if np.sum(overlap) > 100:  # Adjust based on your image resolution
            new_text_polygons.append(poly)

    # Draw polygons on both images
    for poly in new_text_polygons:
        cv2.polylines(before, [np.array(poly, dtype=np.int32)], isClosed=True, color=(36,255,12), thickness=2)
        cv2.polylines(after, [np.array(poly, dtype=np.int32)], isClosed=True, color=(36,255,12), thickness=2)

    # Save or display results
    cv2.imwrite("before_marked.jpg", before)
    cv2.imwrite("after_marked.jpg", after)
    cv2.imshow("Before with New Text Boxes", before)
    cv2.imshow("After with New Text Boxes", after)
    cv2.waitKey(0)
    cv2.destroyAllWindows()

# Run the function
detect_new_text("before.jpg", "after.jpg")

Key Improvements

  • CRAFT Text Detection: Built to handle arbitrary text shapes (rotated, curved, distorted) by predicting character-level regions and grouping them into text blocks.
  • Diff Mask Filtering: Ensures we only mark text that’s actually new (filters out existing text with minor pixel changes).
  • Polygonal Boxes: Draws shapes that perfectly fit curved/rotated text, unlike rigid axis-aligned rectangles.

Alternative: Lightweight EAST Detector

If you need faster performance (and can tolerate slightly less accuracy for curved text), use OpenCV’s EAST detector:

# EAST setup (requires pre-trained weights: frozen_east_text_detection.pb)
net = cv2.dnn.readNet("frozen_east_text_detection.pb")
# Add EAST detection logic here, then filter detected boxes against the diff mask same as above

Troubleshooting Tips

  • Adjust the np.sum(overlap) > 100 threshold: Higher values mean stricter filtering (fewer false positives).
  • For small text, blur the diff mask first: thresh = cv2.GaussianBlur(thresh, (3,3), 0) to reduce noise.
  • If you have GPU access, set cuda=True in the Craft initialization for faster detection.

内容的提问来源于stack exchange,提问作者Mahmudur Rahman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 22:42:42