You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过Google Cloud Vision Document OCR API保留文本布局格式?

Preserving Text Layout with Google Cloud Vision Document OCR API

Great question—you’re right that calling getFullTextAnnotation().getText() dumps you a flat, unformatted string. But the API absolutely gives you the tools to reconstruct the original layout; you just need to work with the hierarchical structure of FullTextAnnotation instead of grabbing the raw text directly.

Here’s how to make it happen:

  • Tap into the nested structure of FullTextAnnotation
    This object isn’t just a single text blob—it’s a layered hierarchy: Page → Block → Paragraph → Word → Symbol, paired with Break elements that explicitly mark line breaks, spaces, tab stops, and other whitespace. Each element also includes bounding box coordinates if you need to handle complex layouts like multi-column text.

    When iterating through content, check for Break objects after each word. The Break.Type enum will tell you exactly what kind of whitespace to add (line break, space, etc.) to match the original document’s formatting.

  • Sample Java code to rebuild formatted text
    Instead of using getText(), loop through the structure and build your text with layout intact:

    FullTextAnnotation fullText = response.getFullTextAnnotation();
    StringBuilder formattedText = new StringBuilder();
    
    for (Page page : fullText.getPagesList()) {
      for (Block block : page.getBlocksList()) {
        for (Paragraph para : block.getParagraphsList()) {
          for (Word word : para.getWordsList()) {
            // Append the word's characters
            for (Symbol symbol : word.getSymbolsList()) {
              formattedText.append(symbol.getText());
            }
            // Handle breaks after the word
            if (!word.getBreakList().isEmpty()) {
              Break breakObj = word.getBreakList().get(0);
              switch (breakObj.getType()) {
                case LINE_BREAK:
                  formattedText.append("\n");
                  break;
                case SPACE:
                  formattedText.append(" ");
                  break;
                // Add cases for TAB, EOL_SURE_SPACE, etc., if needed
                default:
                  formattedText.append(" ");
              }
            }
          }
        }
      }
    }
    
    String layoutPreservedText = formattedText.toString();
    
  • Use structured output formats (optional)
    If you prefer working with serialized data, you can request the OCR output in JSON or XML format via API request parameters. These formats include all hierarchical layout details, which you can parse to build formatted text outside the client library.

This approach lets you retain line breaks, spacing, and basic paragraph structure that the flat getText() call discards.

内容的提问来源于stack exchange,提问作者alexanoid

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:00:41