Google Vision OCR无法识别小文本及倒置图像问题咨询
Hey there, let's fix those two common Google Vision OCR problems you're running into—here are practical, code-ready solutions tailored to your Camera2 + Vision setup:
1. Fixing Inverted Image Recognition
Camera2 captures images based on the sensor's native orientation, which often doesn't match your device's display orientation. This causes the OCR to process upside-down or rotated text, leading to failed recognition.
How to Correct Image Orientation:
- Calculate the required rotation angle by matching the device's display rotation to the camera sensor's orientation.
- Rotate the captured bitmap before passing it to the
Frameobject.
Code Modification:
First, add a helper method to get the correct rotation angle:
private int getCorrectRotationAngle(CameraCharacteristics characteristics, int displayRotation) { int sensorOrientation = characteristics.get(CameraCharacteristics.SENSOR_ORIENTATION); int rotationCompensation; switch (displayRotation) { case Surface.ROTATION_0: rotationCompensation = sensorOrientation; break; case Surface.ROTATION_90: rotationCompensation = (sensorOrientation + 270) % 360; break; case Surface.ROTATION_180: rotationCompensation = (sensorOrientation + 180) % 360; break; case Surface.ROTATION_270: rotationCompensation = (sensorOrientation + 90) % 360; break; default: rotationCompensation = sensorOrientation; break; } return rotationCompensation; }
Then rotate your bitmap before creating the Frame:
// Get the rotation angle (pass CameraCharacteristics from your Camera2 setup) int rotationAngle = getCorrectRotationAngle(cameraCharacteristics, getWindowManager().getDefaultDisplay().getRotation()); // Rotate the bitmap Matrix matrix = new Matrix(); matrix.postRotate(rotationAngle); Bitmap rotatedBitmap = Bitmap.createBitmap(bitmap, 0, 0, bitmap.getWidth(), bitmap.getHeight(), matrix, true); // Now use the rotated bitmap for OCR Frame frame = new Frame.Builder().setBitmap(rotatedBitmap).build(); SparseArray<TextBlock> detectedItems = textRecognizer.detect(frame);
2. Improving Small Text Recognition
Google Vision OCR performs better with high-resolution, clear text. For small text, we can boost recognition by preprocessing the image and adjusting OCR settings.
Key Fixes:
- Scale Up the Image: Enlarge the bitmap (focus on the text region if possible) to increase text size relative to the OCR's detection threshold.
- Image Preprocessing: Convert to grayscale or apply contrast enhancement to make small text more distinct.
- Target ROI (Region of Interest): If you know where the small text is, crop that area and process it separately to avoid wasting resources on irrelevant parts.
Code Examples:
Scaling the Bitmap:
// Scale the bitmap by 2x (adjust factor based on your needs) float scaleFactor = 2.0f; Bitmap scaledBitmap = Bitmap.createScaledBitmap(bitmap, (int)(bitmap.getWidth() * scaleFactor), (int)(bitmap.getHeight() * scaleFactor), true); // Use scaled bitmap for OCR Frame frame = new Frame.Builder().setBitmap(scaledBitmap).build();
Preprocessing for Contrast:
// Convert bitmap to grayscale and enhance contrast Bitmap grayscaleBitmap = Bitmap.createBitmap(bitmap.getWidth(), bitmap.getHeight(), Bitmap.Config.ARGB_8888); Canvas canvas = new Canvas(grayscaleBitmap); Paint paint = new Paint(); ColorMatrix colorMatrix = new ColorMatrix(); colorMatrix.setSaturation(0); // Grayscale ColorMatrixColorFilter filter = new ColorMatrixColorFilter(colorMatrix); paint.setColorFilter(filter); canvas.drawBitmap(bitmap, 0, 0, paint); // Enhance contrast (adjust values as needed) ColorMatrix contrastMatrix = new ColorMatrix(); contrastMatrix.set(new float[]{ 1.5f, 0, 0, 0, -50, // Red channel 0, 1.5f, 0, 0, -50, // Green channel 0, 0, 1.5f, 0, -50, // Blue channel 0, 0, 0, 1, 0 // Alpha channel }); paint.setColorFilter(new ColorMatrixColorFilter(contrastMatrix)); Canvas contrastCanvas = new Canvas(grayscaleBitmap); contrastCanvas.drawBitmap(grayscaleBitmap, 0, 0, paint); // Use preprocessed bitmap for OCR Frame frame = new Frame.Builder().setBitmap(grayscaleBitmap).build();
Combining Both Fixes in Your Existing Code:
Here's how you can integrate both fixes into your original code:
TextRecognizer textRecognizer = TextRecognizerController.getControllerInstance().getTextRecognizerInstance(getContext(), getView()); if (textRecognizer != null && textRecognizer.isOperational()) { // Step 1: Correct image orientation int rotationAngle = getCorrectRotationAngle(cameraCharacteristics, getWindowManager().getDefaultDisplay().getRotation()); Matrix matrix = new Matrix(); matrix.postRotate(rotationAngle); Bitmap rotatedBitmap = Bitmap.createBitmap(bitmap, 0, 0, bitmap.getWidth(), bitmap.getHeight(), matrix, true); // Step 2: Scale and preprocess for small text float scaleFactor = 2.0f; Bitmap scaledBitmap = Bitmap.createScaledBitmap(rotatedBitmap, (int)(rotatedBitmap.getWidth() * scaleFactor), (int)(rotatedBitmap.getHeight() * scaleFactor), true); // Create frame with processed bitmap Frame frame = new Frame.Builder().setBitmap(scaledBitmap).build(); SparseArray<TextBlock> detectedItems = textRecognizer.detect(frame); if (detectedItems != null && detectedItems.size() != 0) { // Your existing processing logic here } }
Additional Tips:
- Ensure your
TextRecognizeris using the latest model by enabling auto-update in the Vision API settings. - If possible, guide users to align small text within the center of the camera frame for better capture quality.
内容的提问来源于stack exchange,提问作者Ragini

