You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Google Cloud OCR,如何在扫描图像PDF中定位并高亮文本段落?

Great question! I’ve tackled this exact scenario before when working with scanned PDFs and Google Cloud OCR—here’s a step-by-step approach to get paragraph-level highlighting working smoothly:

1. Leverage OCR Hierarchy or Group Words into Paragraphs

First off, if you’re only pulling a flat list of words from Google Cloud OCR, you might be missing out on built-in structure. The full OCR response organizes content into blocks → paragraphs → words, and each paragraph comes with its own pre-calculated bounding box. If you can switch to using the paragraphs field directly, that’s the fastest path to getting paragraph positions.

But if you’re stuck with just the word array, no problem—we can group words into paragraphs manually using their bounding box coordinates:

  • Group words into lines: Compare the vertical (y-axis) positions of each word’s bounding box. Words whose y-ranges (top to bottom edges) overlap by a large threshold (like 80% of the word’s height) belong to the same line.
  • Group lines into paragraphs: Look at the vertical gap between the bottom of one line and the top of the next. If the gap is bigger than 1.5x the average line height, that’s a new paragraph. Indentation is another clue—if a line starts significantly further right than the previous one, it’s likely the start of a new paragraph.

2. Match Your Target Paragraph to OCR Text

Once you have structured paragraphs (either from OCR or your manual grouping), you need to find which one matches your target text. Since OCR can have minor errors (typos, extra spaces), skip exact string matches and use these tactics:

  • Use fuzzy string matching (like Python’s fuzzywuzzy library) to calculate similarity scores between your target text and each OCR paragraph. A score above 80-90% is usually a solid match.
  • For longer paragraphs, break both the target and OCR text into 3-word chunks (n-grams) and match overlapping chunks to account for small OCR mistakes.
  • Narrow down candidates first by searching for unique keywords from your target paragraph—this saves time on full fuzzy matches.

3. Calculate the Paragraph’s Bounding Box (If Needed)

If you built the paragraph from individual words, compute the overall bounding box by aggregating the word coordinates:

  • Take the minimum x-coordinate from all words (left edge of the paragraph)
  • Take the minimum y-coordinate from all words (top edge)
  • Take the maximum x-coordinate from all words (right edge)
  • Take the maximum y-coordinate from all words (bottom edge)
    This gives you a single rectangle that perfectly encloses the entire paragraph.

4. Apply the Highlight to the PDF

Scanned PDFs are essentially images wrapped in a PDF container, so you’ll need to overlay a semi-transparent rectangle on top of the page at the paragraph’s bounding box. PyMuPDF (fitz) is a great tool for this—it’s lightweight and easy to use.

Here’s a quick Python example snippet:

import fitz  # PyMuPDF

# Load your scanned PDF
doc = fitz.open("scanned_document.pdf")
target_page = doc[0]  # Adjust to the page with your target paragraph

# Example paragraph bounding box (from OCR or your calculation)
para_bbox = (x_min, y_min, x_max, y_max)

# Create a semi-transparent yellow highlight
highlight = target_page.add_rect_annot(para_bbox)
highlight.set_colors(stroke=None, fill=fitz.utils.getColor("yellow"))
highlight.set_opacity(0.3)  # Tweak transparency for readability
highlight.update()

# Save the modified PDF
doc.save("highlighted_document.pdf")
doc.close()

Pro Tips

  • Handle multi-column layouts: Google Cloud OCR’s blocks field can help identify columns—each block often corresponds to a single column. Process columns separately to maintain proper reading order.
  • Improve OCR accuracy first: If matching fails consistently, pre-process your scanned images (deskew, enhance contrast, remove noise) before running OCR. This reduces typos and makes matching easier.
  • Batch processing: For multiple PDFs, wrap the logic in a loop that processes each page, runs OCR (if not pre-done), matches paragraphs, and applies highlights automatically.

内容的提问来源于stack exchange,提问作者Leitwerk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 09:24:17