如何通过Google Cloud Vision Document OCR API保留文本布局格式?
Great question—you’re right that calling getFullTextAnnotation().getText() dumps you a flat, unformatted string. But the API absolutely gives you the tools to reconstruct the original layout; you just need to work with the hierarchical structure of FullTextAnnotation instead of grabbing the raw text directly.
Here’s how to make it happen:
Tap into the nested structure of
FullTextAnnotation
This object isn’t just a single text blob—it’s a layered hierarchy:Page→Block→Paragraph→Word→Symbol, paired withBreakelements that explicitly mark line breaks, spaces, tab stops, and other whitespace. Each element also includes bounding box coordinates if you need to handle complex layouts like multi-column text.When iterating through content, check for
Breakobjects after each word. TheBreak.Typeenum will tell you exactly what kind of whitespace to add (line break, space, etc.) to match the original document’s formatting.Sample Java code to rebuild formatted text
Instead of usinggetText(), loop through the structure and build your text with layout intact:FullTextAnnotation fullText = response.getFullTextAnnotation(); StringBuilder formattedText = new StringBuilder(); for (Page page : fullText.getPagesList()) { for (Block block : page.getBlocksList()) { for (Paragraph para : block.getParagraphsList()) { for (Word word : para.getWordsList()) { // Append the word's characters for (Symbol symbol : word.getSymbolsList()) { formattedText.append(symbol.getText()); } // Handle breaks after the word if (!word.getBreakList().isEmpty()) { Break breakObj = word.getBreakList().get(0); switch (breakObj.getType()) { case LINE_BREAK: formattedText.append("\n"); break; case SPACE: formattedText.append(" "); break; // Add cases for TAB, EOL_SURE_SPACE, etc., if needed default: formattedText.append(" "); } } } } } } String layoutPreservedText = formattedText.toString();Use structured output formats (optional)
If you prefer working with serialized data, you can request the OCR output in JSON or XML format via API request parameters. These formats include all hierarchical layout details, which you can parse to build formatted text outside the client library.
This approach lets you retain line breaks, spacing, and basic paragraph structure that the flat getText() call discards.
内容的提问来源于stack exchange,提问作者alexanoid

