如何使用ITextSharp读取PDF并验证指定文本是否存在?
Hey there! Let's figure out how to add that string-checking functionality to your iTextSharp PDF reader. You've already got the basics down for reading PDF text—now let's extend it to find and return your target string.
Option 1: Simple Existence Check (Quick & Easy)
If you just need to verify if the string exists and confirm it, you can modify your existing code to check the combined text directly. Here's how:
using iTextSharp.text.pdf; using iTextSharp.text.pdf.parser; using System; using System.Text; namespace PDFReader { internal class Class1 { private static void Main(string[] args) { try { StringBuilder text = new StringBuilder(); string targetString = "SpecifiedStringHere"; // Your target string here using (PdfReader reader = new PdfReader(@"C:\Users\Me\Downloads\TestPDF.pdf")) { for (int i = 1; i <= reader.NumberOfPages; i++) { text.Append(PdfTextExtractor.GetTextFromPage(reader, i)); } } // Check if the string exists and output results if (text.ToString().Contains(targetString)) { Console.WriteLine($"Success! Found the string: *{targetString}*"); } else { Console.WriteLine($"Sorry, the string *{targetString}* was not found in the PDF."); } } catch (Exception e) { Console.WriteLine("The file could not be read:"); Console.WriteLine(e.Message); } Console.Read(); } } }
Option 2: Fine-Grained Matching (With Linq, As You Thought)
If you want to locate exactly where the string appears (including page numbers and position) and use the Linq syntax you mentioned, you'll need to use LocationTextExtractionStrategy to get individual text chunks instead of the full page text. This lets you filter and sort the chunks:
First, add the System.Linq namespace, then modify your code like this:
using iTextSharp.text.pdf; using iTextSharp.text.pdf.parser; using System; using System.Collections.Generic; using System.Linq; namespace PDFReader { internal class Class1 { private static void Main(string[] args) { try { string targetString = "SpecifiedStringHere"; // Your target string here List<TextChunk> allTextChunks = new List<TextChunk>(); using (PdfReader reader = new PdfReader(@"C:\Users\Me\Downloads\TestPDF.pdf")) { for (int i = 1; i <= reader.NumberOfPages; i++) { // Use LocationTextExtractionStrategy to capture individual text chunks var extractionStrategy = new LocationTextExtractionStrategy(); PdfTextExtractor.GetTextFromPage(reader, i, extractionStrategy); allTextChunks.AddRange(extractionStrategy.GetLocationTextChunks()); } } // Filter, sort, and collect matching chunks (matches your Linq idea!) var matchingChunks = allTextChunks .Where(chunk => chunk.Text.Contains(targetString)) .OrderBy(chunk => chunk.Location.Y) // PDF uses bottom-left origin, so Y increases upward .Reverse() // Reverse to get top-to-bottom order .ToList(); // Output the results if (matchingChunks.Any()) { Console.WriteLine($"Found {matchingChunks.Count} occurrences of *{targetString}*:"); foreach (var chunk in matchingChunks) { Console.WriteLine($"Page {chunk.Location.PageNumber}: {chunk.Text}"); } } else { Console.WriteLine($"No occurrences of *{targetString}* found in the PDF."); } } catch (Exception e) { Console.WriteLine("The file could not be read:"); Console.WriteLine(e.Message); } Console.Read(); } } }
A Quick Note on Sorting:
PDFs use a coordinate system where the origin is the bottom-left corner of the page. That means higher Y-values correspond to positions higher up on the page. By using OrderBy(chunk => chunk.Location.Y).Reverse(), we're sorting the matching chunks from top to bottom of the document—just like you'd read it.
内容的提问来源于stack exchange,提问作者Wheelybob

