You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:基于OpenCV与Camera2Basic实现相机预览文本段落矩形框绘制

Solution: Drawing Paragraph-Level Text Bounding Boxes in Camera2Basic with OpenCV

Hey there! Let's fix that issue where you're detecting individual text elements but can't group them into full paragraph boxes like CamScanner does. Here's a step-by-step approach tailored to your setup with Camera2Basic and OpenCV:

1. First: Ensure Stable Detection of Individual Text Boxes

Before merging boxes, make sure you're reliably getting all the small text bounding boxes (per word/line) from OpenCV. Two common methods here:

  • EAST Text Detector: Use OpenCV's DNN module to load the EAST model, process each preview frame, and decode the output scores and geometry to get individual text rectangles.
  • Tesseract OCR: If you're using OCR alongside detection, call getBoundingBoxes() to retrieve coordinates for every recognized text block.

2. Merge Small Boxes into Paragraph-Level Rectangles

The core challenge is grouping adjacent/related text boxes into a single paragraph box. Try these practical methods:

Option A: Use OpenCV's Built-in groupRectangles

This function is designed to cluster overlapping or nearby rectangles—perfect for merging text lines into paragraphs:

// Assume you have a vector<Rect> containing all individual text boxes
vector<Rect> textRects;
vector<Rect> mergedParagraphs;
vector<int> groupWeights;

// Adjust parameters based on your needs:
// - groupThreshold: Minimum number of rects needed to form a group
// - eps: Maximum relative distance between rects to be grouped (0.5 = 50% of rect size)
cv::groupRectangles(textRects, groupWeights, 2, 0.5);

// mergedParagraphs now holds your paragraph-level bounding boxes

Option B: Custom Row-Based Clustering (For More Control)

If you need finer control over how lines are grouped into paragraphs:

  • Step 1: Sort all text boxes by their y-coordinate to separate rows of text.
  • Step 2: For each row, merge boxes that are horizontally close (e.g., spacing less than half the row height).
  • Step 3: Group consecutive rows into paragraphs if their vertical spacing is less than 1.5x the row height (adjust based on typical text spacing).

3. Fix Coordinate Mismatch Between Camera Frame and Screen

Camera2 preview frames are often oriented differently than your screen (e.g., sensor is landscape, screen is portrait). You need to transform your bounding boxes to match the on-screen preview:

  • Get the camera sensor orientation from CameraCharacteristics.SENSOR_ORIENTATION.
  • Calculate the required rotation angle to align the frame with your screen's orientation.
  • Apply the same rotation/scale transformation to your merged paragraph boxes that you apply to the preview frame itself. For example, if you rotate the frame 90 degrees clockwise, use cv::getRotationMatrix2D to compute the transformation and apply it to each rectangle's coordinates.

4. Draw Boxes on the Camera Preview

In Camera2Basic, you have a few options to render the boxes:

  • Draw on the OpenCV Mat: Use Imgproc.rectangle() to draw directly on the processed frame before converting it to a Bitmap for display in the TextureView.
  • Overlay on TextureView: Use a Canvas to draw the transformed boxes on top of the TextureView's surface in the onDraw() method.

Example Code Snippet (Java)

Here's a quick snippet integrating this into the Camera2Basic ImageReader callback:

@Override
public void onImageAvailable(ImageReader reader) {
    Image image = reader.acquireLatestImage();
    if (image == null) return;

    // Convert Camera2 Image to OpenCV Mat
    Mat frame = convertImageToMat(image);
    
    // Rotate frame to match screen orientation
    rotateFrame(frame, sensorOrientation);
    
    // Detect individual text boxes
    List<Rect> textRects = detectTextWithOpenCV(frame);
    
    // Merge into paragraph boxes
    List<Rect> paragraphRects = mergeTextRects(textRects);
    
    // Draw boxes on the frame
    for (Rect rect : paragraphRects) {
        Imgproc.rectangle(frame, rect.tl(), rect.br(), new Scalar(0, 255, 0), 3);
    }
    
    // Convert back to Bitmap and update TextureView
    Bitmap bitmap = convertMatToBitmap(frame);
    runOnUiThread(() -> textureView.setImageBitmap(bitmap));
    
    image.close();
}

Pro Tips for CamScanner-Like Polish

  • Perspective Correction: Once you have the paragraph box, use cv::getPerspectiveTransform to warp the text region into a straight rectangle, just like CamScanner's auto-crop.
  • Performance Optimization: Process downscaled frames for faster detection, then scale the bounding boxes back to the original preview size. Run all detection/merging logic on a background thread to avoid UI lag.

内容的提问来源于stack exchange,提问作者Shankar Kumar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:36:23