Python调用Google Cloud Vision API OCR结果格式化问题
Hey there! Let's fix up your OCR formatting issue. The problem with your current code is that you're returning immediately on the first iteration of the for text in texts loop—so you only ever get the first text annotation, which is why you're missing the rest of the content.
Here's how Google Cloud Vision's TextAnnotation response works:
response.text_annotations[0]contains the full concatenated text from the image, including newline characters where the API detected natural line breaks.- The subsequent entries in
text_annotationsare individual text segments (like words or characters) with position metadata, but you don't need those for just formatting the full readable text.
Here's the revised function that will split the text into lines, add four spaces to the start of each line, and append a newline to each:
def detect_text_uri(uri): """Detects text in the file located in Google Cloud Storage or on the Web. """ client = vision.ImageAnnotatorClient() image = types.Image() image.source.image_uri = uri response = client.text_detection(image=image) # Handle any API errors first if response.error.message: raise Exception( f'{response.error.message}\n' 'For details on error handling, refer to Google Cloud Vision docs.' ) # Check if any text was detected if not response.text_annotations: return "" # Grab the full detected text (includes natural line breaks) full_detected_text = response.text_annotations[0].description # Split into individual lines based on OCR-detected breaks text_lines = full_detected_text.splitlines() # Format each line: 4 leading spaces + line content + newline formatted_text = "".join([f" {line}\n" for line in text_lines]) return formatted_text
Let me break down the key changes:
- Removed the loop that caused an early return, so we no longer truncate the text after the first annotation
- Added basic error handling for API issues and cases where no text was detected
- Used
splitlines()to split the full text using the line breaks the OCR API identified in the image - Used a list comprehension to apply the formatting to every line, then joined them into a single string
This should give you properly formatted text where every line starts with four spaces and ends with a newline, instead of only the first line working as expected.
内容的提问来源于stack exchange,提问作者Tux

