如何用R(含Tesseract导入场景)检索关键词并返回完整语句
Let's walk through this task step by step—from turning your scanned TIFF into usable text to pulling out the exact sentences that contain your target keyword.
Step 1: Install & Load Required Packages
First, grab the tools we need: tesseract for OCR (text extraction from images) and stringr for easy text manipulation:
# Install packages if you haven't already install.packages(c("tesseract", "stringr")) # Load the packages into your R session library(tesseract) library(stringr)
Step 2: Pull Text from Your TIFF File
Use Tesseract to convert the scanned image into raw text. For multi-page TIFFs, you’ll need to loop through each page, but here’s the single-page version to start with:
# Replace with your actual TIFF file path img_path <- "your_scanned_document.tiff" # Run OCR to extract text from the image raw_text <- ocr(img_path)
Step 3: Clean Text & Split into Full Sentences
OCR output often has messy line breaks and extra spaces. Let’s fix that first, then split the text into complete sentences using regex:
# Remove line breaks and replace multiple spaces with a single space clean_text <- str_replace_all(raw_text, "\\n|\\s{2,}", " ") # Split text into sentences (handles . ! ? and ellipses …) sentences <- str_split(clean_text, "(?<=[.!?…])\\s+")[[1]]
Quick note: The regex (?<=[.!?…])\\s+ splits text right after sentence-ending punctuation, so we keep full, intact sentences instead of chopping them up.
Step 4: Retrieve Sentences with Your Target Keyword
Define your keyword, then filter the sentences to only those that include it (we’ll use case-insensitive matching to catch variations like "Experience" or "EXPERIENCE"):
# Replace with your target keyword target_keyword <- "experience" # Filter sentences that contain the keyword matching_sentences <- sentences[grepl(target_keyword, sentences, ignore.case = TRUE)] # Print the results to see your matched sentences print(matching_sentences)
Example Output
Using your sample text, this would return:
Look at the years of experience required and see if that matches the years of experience that the candidate has.
Quick Troubleshooting Tips
- If sentence splitting is wonky: Adjust the regex to account for OCR errors (like missing spaces after punctuation). For example, try
str_split(clean_text, "(?<=[.!?…])(?=\\p{Lu})")to split before uppercase letters that follow sentence endings. - If OCR text is messy: Preprocess your TIFF first (e.g., increase contrast, crop out irrelevant edges) to improve text quality before running OCR.
内容的提问来源于stack exchange,提问作者fallynkay

