You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Python实现阿拉伯文本检测框的分组与重叠分离

Hey there! Let’s break down how to tackle your Arabic text bounding box tasks—merging tiny boxes into meaningful larger ones and separating overlapping ones. Since Arabic is right-to-left (RTL), we’ll need to keep that directionality in mind for some steps, but the core logic is straightforward with a few targeted tweaks.

Step 1: Extract and Format Bounding Box Data

First, you’ll need to pull bounding boxes from your image along with their associated text and confidence scores. Tools like EasyOCR work great for Arabic text since it supports RTL natively. Here’s a quick snippet to get you started:

import easyocr

# Initialize reader for Arabic
reader = easyocr.Reader(['ar'])
# Extract text and boxes from your image
result = reader.readtext('your_arabic_image.jpg')

# Convert the raw bbox format (4 points) to a simpler (x_min, y_min, x_max, y_max) structure
formatted_boxes = []
for item in result:
    bbox_points = item[0]
    x_coords = [p[0] for p in bbox_points]
    y_coords = [p[1] for p in bbox_points]
    x_min, x_max = min(x_coords), max(x_coords)
    y_min, y_max = min(y_coords), max(y_coords)
    # Store box coords, text, and confidence
    formatted_boxes.append( (x_min, y_min, x_max, y_max, item[1], item[2]) )
Step 2: Merge Small Bounding Boxes into Larger Ones

Small boxes usually come from isolated characters, punctuation, or OCR errors. To merge them into valid text blocks:

  1. Define "small": Set an area threshold (adjust based on your image size—e.g., 500 pixels² for standard resolutions).
  2. Group boxes by rows: Cluster boxes that sit on the same horizontal line (using vertical overlap as a metric).
  3. Sort rows for RTL: Since Arabic reads right-to-left, sort each row’s boxes by their x_max (rightmost point) in descending order.
  4. Merge adjacent small boxes: For each tiny box, check nearby larger boxes in the same row. If they’re vertically aligned (high overlap) and horizontally close, merge them into a single box.

Here’s a code snippet implementing this:

# Group boxes into rows (simplified clustering by vertical overlap)
rows = []
# First sort all boxes by their top y-coordinate
sorted_by_y = sorted(formatted_boxes, key=lambda b: b[1])
current_row = [sorted_by_y[0]]

for box in sorted_by_y[1:]:
    prev_box = current_row[-1]
    # Calculate vertical overlap percentage
    y_overlap = min(box[3], prev_box[3]) - max(box[1], prev_box[1])
    avg_height = ( (box[3]-box[1]) + (prev_box[3]-prev_box[1]) ) / 2
    # If boxes share >70% of their height, they're in the same row
    if y_overlap / avg_height > 0.7:
        current_row.append(box)
    else:
        rows.append(current_row)
        current_row = [box]
rows.append(current_row)

# Merge small boxes in each row
merged_boxes = []
area_threshold = 500  # Adjust this based on your image

for row in rows:
    # Sort row in RTL order (rightmost boxes first)
    rtl_sorted = sorted(row, key=lambda b: -b[2])
    i = 0
    while i < len(rtl_sorted):
        current_box = rtl_sorted[i]
        current_area = (current_box[2]-current_box[0]) * (current_box[3]-current_box[1])
        
        if current_area < area_threshold:
            # Look for a nearby large box to merge with
            merged = False
            for j in range(i+1, len(rtl_sorted)):
                target_box = rtl_sorted[j]
                target_area = (target_box[2]-target_box[0]) * (target_box[3]-target_box[1])
                
                if target_area >= area_threshold:
                    # Check vertical alignment and horizontal proximity
                    y_overlap = min(current_box[3], target_box[3]) - max(current_box[1], target_box[1])
                    avg_height = ( (current_box[3]-current_box[1]) + (target_box[3]-target_box[1]) ) / 2
                    # For RTL, small boxes often sit to the left of larger word boxes
                    horizontal_gap = target_box[2] - current_box[0]
                    
                    if y_overlap / avg_height > 0.8 and horizontal_gap < 20:
                        # Merge the two boxes
                        new_x_min = min(current_box[0], target_box[0])
                        new_x_max = max(current_box[2], target_box[2])
                        new_y_min = min(current_box[1], target_box[1])
                        new_y_max = max(current_box[3], target_box[3])
                        # Combine text (note: RTL order might need swapping—adjust based on text content)
                        merged_text = target_box[4] + current_box[4]
                        merged_conf = (current_box[5] + target_box[5]) / 2
                        
                        # Replace target box with merged result, skip current box
                        rtl_sorted[j] = (new_x_min, new_y_min, new_x_max, new_y_max, merged_text, merged_conf)
                        merged = True
                        break
            if merged:
                i += 1
            else:
                # Keep the small box if no merge candidate is found
                merged_boxes.append(current_box)
                i += 1
        else:
            merged_boxes.append(current_box)
            i += 1
Step 3: Separate Overlapping Bounding Boxes

Overlaps usually happen when OCR confuses adjacent words or misaligns rows. Here’s how to fix them:

  • Same-row overlaps: Use IoU (Intersection over Union) to detect overlapping boxes. If IoU exceeds a threshold (e.g., 0.2), split the boxes at the midpoint of their overlapping horizontal range (since RTL, the right box is the earlier word).
  • Cross-row overlaps: Re-adjust your row clustering threshold (reduce vertical overlap to 50-60%) or use a vertical projection method (count pixel density per row to find natural gaps between lines).

Here’s a snippet for fixing same-row overlaps:

final_boxes = []
# Re-group merged boxes into rows (repeat the row clustering code above if needed)
# ... (row grouping code here)

for row in rows:
    rtl_sorted = sorted(row, key=lambda b: -b[2])
    i = 0
    while i < len(rtl_sorted)-1:
        box1 = rtl_sorted[i]
        box2 = rtl_sorted[i+1]
        
        # Calculate intersection area
        x_overlap = min(box1[2], box2[2]) - max(box1[0], box2[0])
        y_overlap = min(box1[3], box2[3]) - max(box1[1], box2[1])
        
        if x_overlap > 0 and y_overlap > 0:
            intersection = x_overlap * y_overlap
            area1 = (box1[2]-box1[0])*(box1[3]-box1[1])
            area2 = (box2[2]-box2[0])*(box2[3]-box2[1])
            union = area1 + area2 - intersection
            iou = intersection / union
            
            if iou > 0.2:
                # Split at the midpoint of the overlapping x range
                overlap_start = max(box1[0], box2[0])
                overlap_end = min(box1[2], box2[2])
                split_x = (overlap_start + overlap_end) / 2
                
                # Adjust boxes to remove overlap
                fixed_box1 = (box1[0], box1[1], split_x, box1[3], box1[4], box1[5])
                fixed_box2 = (split_x, box2[1], box2[2], box2[3], box2[4], box2[5])
                
                rtl_sorted[i] = fixed_box1
                rtl_sorted[i+1] = fixed_box2
        
        final_boxes.append(rtl_sorted[i])
        i += 1
final_boxes.append(rtl_sorted[-1])
Bonus Tips for Arabic Text
  • Filter low-confidence boxes: Before merging, discard boxes with confidence scores < 0.5—these are often OCR noise that don’t belong to valid text.
  • Use Arabic-specific OCR models: Models fine-tuned on Arabic text (like AraBERT-based OCR tools) generate more accurate bounding boxes, reducing the need for manual fixes.
  • Tweak thresholds: All the numbers (area, overlap, gap) depend on your image’s resolution and text size—experiment with small adjustments to get the best results.

Hope these steps help you nail the desired bounding box output!

内容的提问来源于stack exchange,提问作者Hedy Krid

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:36:34