如何定位图像中新增的普通/弯曲/旋转文本区域并绘制包围框
Hey there! The problem with your current code is that it uses basic contour detection on pixel differences, which falls flat for rotated or curved text—those often have scattered diff pixels, and axis-aligned rectangles don’t fit their actual shape. Let’s solve this by pairing text detection models (built for arbitrary text shapes) with your existing image diff logic to zero in on exactly the new text regions.
We’ll use a pre-trained text detection model (CRAFT, which excels at curved/rotated text) alongside image difference filtering. This ensures we only target newly added text, not existing content or non-text differences, and draw precise bounding polygons instead of rigid rectangles.
Step 1: Install Dependencies
First, grab the required packages:
pip install opencv-python numpy scikit-image craft-text-detector
Step 2: Full Implementation Code
This code combines diff analysis with CRAFT to detect and draw accurate bounding shapes for all text types:
import cv2 import numpy as np from skimage.metrics import structural_similarity as compare_ssim from craft_text_detector import Craft def detect_new_text(before_img_path, after_img_path): # Load and preprocess images before = cv2.imread(before_img_path) after = cv2.imread(after_img_path) before_gray = cv2.cvtColor(before, cv2.COLOR_BGR2GRAY) after_gray = cv2.cvtColor(after, cv2.COLOR_BGR2GRAY) # Generate difference mask using SSIM (_, diff) = compare_ssim(before_gray, after_gray, full=True) diff = (diff * 255).astype("uint8") thresh = cv2.threshold(diff, 0, 255, cv2.THRESH_BINARY_INV | cv2.THRESH_OTSU)[1] # Initialize CRAFT text detector (handles curved/rotated text) craft = Craft(output_dir=None, crop_type="poly", cuda=False) text_preds = craft.detect_text(after) all_text_polygons = text_preds["polygons"] # Filter polygons to only those in the difference region (new text) new_text_polygons = [] for poly in all_text_polygons: # Create mask for current text polygon poly_mask = np.zeros_like(thresh) cv2.fillPoly(poly_mask, [np.array(poly, dtype=np.int32)], 255) # Check if polygon overlaps significantly with diff mask overlap = cv2.bitwise_and(thresh, poly_mask) if np.sum(overlap) > 100: # Adjust based on your image resolution new_text_polygons.append(poly) # Draw polygons on both images for poly in new_text_polygons: cv2.polylines(before, [np.array(poly, dtype=np.int32)], isClosed=True, color=(36,255,12), thickness=2) cv2.polylines(after, [np.array(poly, dtype=np.int32)], isClosed=True, color=(36,255,12), thickness=2) # Save or display results cv2.imwrite("before_marked.jpg", before) cv2.imwrite("after_marked.jpg", after) cv2.imshow("Before with New Text Boxes", before) cv2.imshow("After with New Text Boxes", after) cv2.waitKey(0) cv2.destroyAllWindows() # Run the function detect_new_text("before.jpg", "after.jpg")
Key Improvements
- CRAFT Text Detection: Built to handle arbitrary text shapes (rotated, curved, distorted) by predicting character-level regions and grouping them into text blocks.
- Diff Mask Filtering: Ensures we only mark text that’s actually new (filters out existing text with minor pixel changes).
- Polygonal Boxes: Draws shapes that perfectly fit curved/rotated text, unlike rigid axis-aligned rectangles.
Alternative: Lightweight EAST Detector
If you need faster performance (and can tolerate slightly less accuracy for curved text), use OpenCV’s EAST detector:
# EAST setup (requires pre-trained weights: frozen_east_text_detection.pb) net = cv2.dnn.readNet("frozen_east_text_detection.pb") # Add EAST detection logic here, then filter detected boxes against the diff mask same as above
Troubleshooting Tips
- Adjust the
np.sum(overlap) > 100threshold: Higher values mean stricter filtering (fewer false positives). - For small text, blur the diff mask first:
thresh = cv2.GaussianBlur(thresh, (3,3), 0)to reduce noise. - If you have GPU access, set
cuda=Truein the Craft initialization for faster detection.
内容的提问来源于stack exchange,提问作者Mahmudur Rahman

