You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Android Jetpack Compose下CameraX+ML Kit实现ROI文本识别的技术问询

CameraX + ML Kit 实现ROI区域文本识别解决方案

问题描述

我正在开发基于Jetpack Compose的Android应用,采用CameraX实现相机预览、ML Kit进行文本识别。目前文本识别会处理整个预览画面,希望仅针对在Composable中定义为Rect的ROI(感兴趣区域)进行识别。

现有代码片段

已创建TextRecognitionAnalyzer类:

class TextRecognitionAnalyzer(
    private val mrzRect: Rect,
    private val screenSize: IntSize,
    private val onTextDetected: (Text) -> Unit
) : ImageAnalysis.Analyzer {
    // Analyzer implementation...
}

当前尝试计算裁剪区域,但存在困惑:

val cropLeft = // calculate crop left based on ROI and scaling
val cropTop = // calculate crop top based on ROI and scaling
val cropRight = // calculate crop right based on ROI and scaling
val cropBottom = // calculate crop bottom based on ROI and scaling

// Use the cropped area for text recognition...

目前遇到识别结果超出ROI范围的问题,以下是针对性解决方案:


1. 基于ROI准确计算图像裁剪坐标

核心是完成屏幕坐标系到CameraX图像坐标系的转换,需处理图像旋转和缩放比例两个关键因素:

override fun analyze(imageProxy: ImageProxy) {
    val mediaImage = imageProxy.image ?: return
    val imageSize = IntSize(mediaImage.width, mediaImage.height)
    val rotationDegrees = imageProxy.imageInfo.rotationDegrees

    // 计算屏幕到图像的缩放比例(适配旋转后的尺寸匹配)
    val (scaleX, scaleY) = when (rotationDegrees) {
        0, 180 -> {
            imageSize.width / screenSize.width.toFloat() to imageSize.height / screenSize.height.toFloat()
        }
        90, 270 -> {
            // 旋转90/270度后,图像宽高与屏幕宽高互换
            imageSize.height / screenSize.width.toFloat() to imageSize.width / screenSize.height.toFloat()
        }
        else -> 1f to 1f
    }

    // 将屏幕ROI转换为图像坐标系下的Rect(处理旋转映射)
    val mappedRect = when (rotationDegrees) {
        0 -> Rect(
            (mrzRect.left * scaleX).toInt(),
            (mrzRect.top * scaleY).toInt(),
            (mrzRect.right * scaleX).toInt(),
            (mrzRect.bottom * scaleY).toInt()
        )
        90 -> Rect(
            // 旋转90度:屏幕x → 图像y轴反向,屏幕y → 图像x轴
            ((screenSize.height - mrzRect.bottom) * scaleX).toInt(),
            (mrzRect.left * scaleY).toInt(),
            ((screenSize.height - mrzRect.top) * scaleX).toInt(),
            (mrzRect.right * scaleY).toInt()
        )
        180 -> Rect(
            ((screenSize.width - mrzRect.right) * scaleX).toInt(),
            ((screenSize.height - mrzRect.bottom) * scaleY).toInt(),
            ((screenSize.width - mrzRect.left) * scaleX).toInt(),
            ((screenSize.height - mrzRect.top) * scaleY).toInt()
        )
        270 -> Rect(
            (mrzRect.top * scaleX).toInt(),
            ((screenSize.width - mrzRect.right) * scaleY).toInt(),
            (mrzRect.bottom * scaleX).toInt(),
            ((screenSize.width - mrzRect.left) * scaleY).toInt()
        )
        else -> mrzRect
    }

    // 确保裁剪区域不超出图像边界,避免崩溃
    val cropLeft = mappedRect.left.coerceIn(0, imageSize.width)
    val cropTop = mappedRect.top.coerceIn(0, imageSize.height)
    val cropRight = mappedRect.right.coerceIn(0, imageSize.width)
    val cropBottom = mappedRect.bottom.coerceIn(0, imageSize.height)

    // 后续裁剪/识别逻辑...
    imageProxy.close()
}

2. 确保文本识别仅在ROI内运行

有两种可靠实现方式,按需选择:

方式一:裁剪图像后再识别(性能更优)

将CameraX输出的图像裁剪到ROI区域,再传给ML Kit处理,减少识别范围:

// 工具方法:将ImageProxy转换为Bitmap并裁剪
private fun cropImageToRoi(imageProxy: ImageProxy, cropRect: Rect): Bitmap {
    val bitmap = imageProxy.toBitmap()
    return Bitmap.createBitmap(
        bitmap,
        cropRect.left,
        cropRect.top,
        cropRect.width(),
        cropRect.height()
    )
}

// 在analyze方法中使用
val croppedBitmap = cropImageToRoi(imageProxy, mappedRect)
val inputImage = InputImage.fromBitmap(croppedBitmap, rotationDegrees)

TextRecognition.getClient().process(inputImage)
    .addOnSuccessListener { text ->
        onTextDetected(text)
    }
    .addOnCompleteListener {
        imageProxy.close()
    }

方式二:过滤识别结果(无需裁剪图像)

如果不希望裁剪图像,可在ML Kit返回结果后,过滤掉ROI外的文本块:

val inputImage = InputImage.fromMediaImage(mediaImage, rotationDegrees)

TextRecognition.getClient().process(inputImage)
    .addOnSuccessListener { text ->
        // 过滤出完全在ROI内的文本块
        val filteredBlocks = text.blocks.filter { block ->
            block.boundingBox?.let { blockRect ->
                mappedRect.contains(blockRect)
            } ?: false
        }
        // 构建过滤后的Text对象返回
        val filteredText = Text.create(text.text, filteredBlocks, text.language)
        onTextDetected(filteredText)
    }
    .addOnCompleteListener {
        imageProxy.close()
    }

3. 最佳实践与避坑要点

  • 必须处理图像旋转:CameraX输出的图像旋转角度由imageProxy.imageInfo.rotationDegrees决定,不能硬编码(不同设备传感器方向可能不同)
  • 缩放比例适配旋转:旋转90/270度后,图像宽高与屏幕宽高互换,缩放因子需对应调整
  • 边界校验不可少:计算出的裁剪坐标必须通过coerceIn限制在图像范围内,避免数组越界崩溃
  • 优先选择裁剪图像:裁剪后能减少ML Kit的计算量,提升识别速度和响应性
  • ROI与预览同步:不要直接用屏幕尺寸计算ROI,建议基于CameraX的PreviewView实际尺寸(避免预览有黑边导致坐标偏移)
  • 多设备测试:不同品牌设备的相机传感器方向可能有差异,需在多种设备上验证坐标转换准确性

内容的提问来源于stack exchange,提问作者Hocine Elhadj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 18:35:09