You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python调用Google Cloud Vision API OCR结果格式化问题

Hey there! Let's fix up your OCR formatting issue. The problem with your current code is that you're returning immediately on the first iteration of the for text in texts loop—so you only ever get the first text annotation, which is why you're missing the rest of the content.

Here's how Google Cloud Vision's TextAnnotation response works:

  • response.text_annotations[0] contains the full concatenated text from the image, including newline characters where the API detected natural line breaks.
  • The subsequent entries in text_annotations are individual text segments (like words or characters) with position metadata, but you don't need those for just formatting the full readable text.

Here's the revised function that will split the text into lines, add four spaces to the start of each line, and append a newline to each:

def detect_text_uri(uri):
    """Detects text in the file located in Google Cloud Storage or on the Web. """
    client = vision.ImageAnnotatorClient()
    image = types.Image()
    image.source.image_uri = uri
    response = client.text_detection(image=image)

    # Handle any API errors first
    if response.error.message:
        raise Exception(
            f'{response.error.message}\n'
            'For details on error handling, refer to Google Cloud Vision docs.'
        )

    # Check if any text was detected
    if not response.text_annotations:
        return ""

    # Grab the full detected text (includes natural line breaks)
    full_detected_text = response.text_annotations[0].description
    # Split into individual lines based on OCR-detected breaks
    text_lines = full_detected_text.splitlines()
    # Format each line: 4 leading spaces + line content + newline
    formatted_text = "".join([f"    {line}\n" for line in text_lines])

    return formatted_text

Let me break down the key changes:

  • Removed the loop that caused an early return, so we no longer truncate the text after the first annotation
  • Added basic error handling for API issues and cases where no text was detected
  • Used splitlines() to split the full text using the line breaks the OCR API identified in the image
  • Used a list comprehension to apply the formatting to every line, then joined them into a single string

This should give you properly formatted text where every line starts with four spaces and ends with a newline, instead of only the first line working as expected.

内容的提问来源于stack exchange,提问作者Tux

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:10:50