You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用R(含Tesseract导入场景)检索关键词并返回完整语句

Extract Full Sentences with Target Keywords from OCR'd TIFF in R

Let's walk through this task step by step—from turning your scanned TIFF into usable text to pulling out the exact sentences that contain your target keyword.

Step 1: Install & Load Required Packages

First, grab the tools we need: tesseract for OCR (text extraction from images) and stringr for easy text manipulation:

# Install packages if you haven't already
install.packages(c("tesseract", "stringr"))

# Load the packages into your R session
library(tesseract)
library(stringr)

Step 2: Pull Text from Your TIFF File

Use Tesseract to convert the scanned image into raw text. For multi-page TIFFs, you’ll need to loop through each page, but here’s the single-page version to start with:

# Replace with your actual TIFF file path
img_path <- "your_scanned_document.tiff"

# Run OCR to extract text from the image
raw_text <- ocr(img_path)

Step 3: Clean Text & Split into Full Sentences

OCR output often has messy line breaks and extra spaces. Let’s fix that first, then split the text into complete sentences using regex:

# Remove line breaks and replace multiple spaces with a single space
clean_text <- str_replace_all(raw_text, "\\n|\\s{2,}", " ")

# Split text into sentences (handles . ! ? and ellipses …)
sentences <- str_split(clean_text, "(?<=[.!?…])\\s+")[[1]]

Quick note: The regex (?<=[.!?…])\\s+ splits text right after sentence-ending punctuation, so we keep full, intact sentences instead of chopping them up.

Step 4: Retrieve Sentences with Your Target Keyword

Define your keyword, then filter the sentences to only those that include it (we’ll use case-insensitive matching to catch variations like "Experience" or "EXPERIENCE"):

# Replace with your target keyword
target_keyword <- "experience"

# Filter sentences that contain the keyword
matching_sentences <- sentences[grepl(target_keyword, sentences, ignore.case = TRUE)]

# Print the results to see your matched sentences
print(matching_sentences)

Example Output

Using your sample text, this would return:

Look at the years of experience required and see if that matches the years of experience that the candidate has.

Quick Troubleshooting Tips

  • If sentence splitting is wonky: Adjust the regex to account for OCR errors (like missing spaces after punctuation). For example, try str_split(clean_text, "(?<=[.!?…])(?=\\p{Lu})") to split before uppercase letters that follow sentence endings.
  • If OCR text is messy: Preprocess your TIFF first (e.g., increase contrast, crop out irrelevant edges) to improve text quality before running OCR.

内容的提问来源于stack exchange,提问作者fallynkay

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 04:20:14