基于Python实现阿拉伯文本检测框的分组与重叠分离
Hey there! Let’s break down how to tackle your Arabic text bounding box tasks—merging tiny boxes into meaningful larger ones and separating overlapping ones. Since Arabic is right-to-left (RTL), we’ll need to keep that directionality in mind for some steps, but the core logic is straightforward with a few targeted tweaks.
First, you’ll need to pull bounding boxes from your image along with their associated text and confidence scores. Tools like EasyOCR work great for Arabic text since it supports RTL natively. Here’s a quick snippet to get you started:
import easyocr # Initialize reader for Arabic reader = easyocr.Reader(['ar']) # Extract text and boxes from your image result = reader.readtext('your_arabic_image.jpg') # Convert the raw bbox format (4 points) to a simpler (x_min, y_min, x_max, y_max) structure formatted_boxes = [] for item in result: bbox_points = item[0] x_coords = [p[0] for p in bbox_points] y_coords = [p[1] for p in bbox_points] x_min, x_max = min(x_coords), max(x_coords) y_min, y_max = min(y_coords), max(y_coords) # Store box coords, text, and confidence formatted_boxes.append( (x_min, y_min, x_max, y_max, item[1], item[2]) )
Small boxes usually come from isolated characters, punctuation, or OCR errors. To merge them into valid text blocks:
- Define "small": Set an area threshold (adjust based on your image size—e.g., 500 pixels² for standard resolutions).
- Group boxes by rows: Cluster boxes that sit on the same horizontal line (using vertical overlap as a metric).
- Sort rows for RTL: Since Arabic reads right-to-left, sort each row’s boxes by their
x_max(rightmost point) in descending order. - Merge adjacent small boxes: For each tiny box, check nearby larger boxes in the same row. If they’re vertically aligned (high overlap) and horizontally close, merge them into a single box.
Here’s a code snippet implementing this:
# Group boxes into rows (simplified clustering by vertical overlap) rows = [] # First sort all boxes by their top y-coordinate sorted_by_y = sorted(formatted_boxes, key=lambda b: b[1]) current_row = [sorted_by_y[0]] for box in sorted_by_y[1:]: prev_box = current_row[-1] # Calculate vertical overlap percentage y_overlap = min(box[3], prev_box[3]) - max(box[1], prev_box[1]) avg_height = ( (box[3]-box[1]) + (prev_box[3]-prev_box[1]) ) / 2 # If boxes share >70% of their height, they're in the same row if y_overlap / avg_height > 0.7: current_row.append(box) else: rows.append(current_row) current_row = [box] rows.append(current_row) # Merge small boxes in each row merged_boxes = [] area_threshold = 500 # Adjust this based on your image for row in rows: # Sort row in RTL order (rightmost boxes first) rtl_sorted = sorted(row, key=lambda b: -b[2]) i = 0 while i < len(rtl_sorted): current_box = rtl_sorted[i] current_area = (current_box[2]-current_box[0]) * (current_box[3]-current_box[1]) if current_area < area_threshold: # Look for a nearby large box to merge with merged = False for j in range(i+1, len(rtl_sorted)): target_box = rtl_sorted[j] target_area = (target_box[2]-target_box[0]) * (target_box[3]-target_box[1]) if target_area >= area_threshold: # Check vertical alignment and horizontal proximity y_overlap = min(current_box[3], target_box[3]) - max(current_box[1], target_box[1]) avg_height = ( (current_box[3]-current_box[1]) + (target_box[3]-target_box[1]) ) / 2 # For RTL, small boxes often sit to the left of larger word boxes horizontal_gap = target_box[2] - current_box[0] if y_overlap / avg_height > 0.8 and horizontal_gap < 20: # Merge the two boxes new_x_min = min(current_box[0], target_box[0]) new_x_max = max(current_box[2], target_box[2]) new_y_min = min(current_box[1], target_box[1]) new_y_max = max(current_box[3], target_box[3]) # Combine text (note: RTL order might need swapping—adjust based on text content) merged_text = target_box[4] + current_box[4] merged_conf = (current_box[5] + target_box[5]) / 2 # Replace target box with merged result, skip current box rtl_sorted[j] = (new_x_min, new_y_min, new_x_max, new_y_max, merged_text, merged_conf) merged = True break if merged: i += 1 else: # Keep the small box if no merge candidate is found merged_boxes.append(current_box) i += 1 else: merged_boxes.append(current_box) i += 1
Overlaps usually happen when OCR confuses adjacent words or misaligns rows. Here’s how to fix them:
- Same-row overlaps: Use IoU (Intersection over Union) to detect overlapping boxes. If IoU exceeds a threshold (e.g., 0.2), split the boxes at the midpoint of their overlapping horizontal range (since RTL, the right box is the earlier word).
- Cross-row overlaps: Re-adjust your row clustering threshold (reduce vertical overlap to 50-60%) or use a vertical projection method (count pixel density per row to find natural gaps between lines).
Here’s a snippet for fixing same-row overlaps:
final_boxes = [] # Re-group merged boxes into rows (repeat the row clustering code above if needed) # ... (row grouping code here) for row in rows: rtl_sorted = sorted(row, key=lambda b: -b[2]) i = 0 while i < len(rtl_sorted)-1: box1 = rtl_sorted[i] box2 = rtl_sorted[i+1] # Calculate intersection area x_overlap = min(box1[2], box2[2]) - max(box1[0], box2[0]) y_overlap = min(box1[3], box2[3]) - max(box1[1], box2[1]) if x_overlap > 0 and y_overlap > 0: intersection = x_overlap * y_overlap area1 = (box1[2]-box1[0])*(box1[3]-box1[1]) area2 = (box2[2]-box2[0])*(box2[3]-box2[1]) union = area1 + area2 - intersection iou = intersection / union if iou > 0.2: # Split at the midpoint of the overlapping x range overlap_start = max(box1[0], box2[0]) overlap_end = min(box1[2], box2[2]) split_x = (overlap_start + overlap_end) / 2 # Adjust boxes to remove overlap fixed_box1 = (box1[0], box1[1], split_x, box1[3], box1[4], box1[5]) fixed_box2 = (split_x, box2[1], box2[2], box2[3], box2[4], box2[5]) rtl_sorted[i] = fixed_box1 rtl_sorted[i+1] = fixed_box2 final_boxes.append(rtl_sorted[i]) i += 1 final_boxes.append(rtl_sorted[-1])
- Filter low-confidence boxes: Before merging, discard boxes with confidence scores < 0.5—these are often OCR noise that don’t belong to valid text.
- Use Arabic-specific OCR models: Models fine-tuned on Arabic text (like AraBERT-based OCR tools) generate more accurate bounding boxes, reducing the need for manual fixes.
- Tweak thresholds: All the numbers (area, overlap, gap) depend on your image’s resolution and text size—experiment with small adjustments to get the best results.
Hope these steps help you nail the desired bounding box output!
内容的提问来源于stack exchange,提问作者Hedy Krid

