能否通过Google Vision API选中特定文本?相机扫描应用开发咨询
Absolutely, you can totally implement touch-based focus to target only the "Hello World" text in your camera feed! Let’s break down how to make this work step by step:
1. Capture Touch Coordinates
- First, when the user taps the screen, grab the exact x/y coordinates of the touch event. On Android, use an
OnTouchListenerto pull values fromMotionEvent; on iOS, aUITapGestureRecognizerwill give you the tap point relative to the camera preview view. - Convert these screen coordinates to match the camera's actual frame (you’ll need to adjust for preview scaling, aspect ratio, and device orientation since the preview might not match the raw camera output exactly).
2. Crop the Camera Frame to the Touch Area
- Once you have the adjusted coordinates, define a small bounding box around the tap (like a 200x200 pixel square centered on the tap point). This way, you’re only sending a cropped snippet of the camera frame to the Vision API instead of the entire image—greatly reducing irrelevant text detection.
- For example, on Android, you can create a cropped
Bitmapfrom the full camera frame before passing it to theTextRecognizer.
3. Filter Results for "Hello World"
- Even with cropping, the API might pick up small bits of other text. Add a post-processing step to filter results:
- Iterate through the
TextBlockobjects returned by the Vision API. - Check if any block contains the exact string "Hello World" (or use a case-insensitive match if needed).
- If found, extract your desired details; if not, you can either ignore the result or prompt the user to tap closer to the target text.
- Iterate through the
4. Optional: Add Visual Feedback
- To make the experience intuitive, draw a rectangle on the preview screen at the tap location. This lets the user see exactly which area is being scanned, helping them align their tap perfectly with "Hello World".
Example Snippet (Android)
Here’s a quick code example to illustrate the touch capture and cropping flow:
cameraPreview.setOnTouchListener { _, event -> if (event.action == MotionEvent.ACTION_UP) { val tapX = event.x val tapY = event.y // Convert screen tap coordinates to match camera preview dimensions val scaledX = tapX * cameraPreview.width / cameraPreview.measuredWidth val scaledY = tapY * cameraPreview.height / cameraPreview.measuredHeight // Define a 200x200 crop box centered on the tap val cropSize = 100 val cropLeft = max(0, scaledX - cropSize).toInt() val cropTop = max(0, scaledY - cropSize).toInt() val cropRight = min(cameraPreview.width, scaledX + cropSize).toInt() val cropBottom = min(cameraPreview.height, scaledY + cropSize).toInt() // Crop the original camera frame bitmap val croppedBitmap = Bitmap.createBitmap(originalCameraFrame, cropLeft, cropTop, cropRight - cropLeft, cropBottom - cropTop) // Send cropped bitmap to Vision API for processing scanTargetText(croppedBitmap) } true } private fun scanTargetText(bitmap: Bitmap) { val textRecognizer = TextRecognizer.Builder(context).build() val imageFrame = Frame.Builder().setBitmap(bitmap).build() val recognitionResult = textRecognizer.detect(imageFrame) for (textBlock in recognitionResult.textBlocks) { if (textBlock.text.trim().equals("Hello World", ignoreCase = true)) { // Target text found! Handle your detail extraction here Log.d("TextScanner", "Found Hello World! Full block text: ${textBlock.text}") break } } }
内容的提问来源于stack exchange,提问作者Syed Ammar
相关产品推荐
相关产品推荐

