You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用ITextSharp读取PDF并验证指定文本是否存在?

How to Check for Specific Strings in a PDF with iTextSharp

Hey there! Let's figure out how to add that string-checking functionality to your iTextSharp PDF reader. You've already got the basics down for reading PDF text—now let's extend it to find and return your target string.

Option 1: Simple Existence Check (Quick & Easy)

If you just need to verify if the string exists and confirm it, you can modify your existing code to check the combined text directly. Here's how:

using iTextSharp.text.pdf;
using iTextSharp.text.pdf.parser;
using System;
using System.Text;

namespace PDFReader
{
    internal class Class1
    {
        private static void Main(string[] args)
        {
            try
            {
                StringBuilder text = new StringBuilder();
                string targetString = "SpecifiedStringHere"; // Your target string here

                using (PdfReader reader = new PdfReader(@"C:\Users\Me\Downloads\TestPDF.pdf"))
                {
                    for (int i = 1; i <= reader.NumberOfPages; i++)
                    {
                        text.Append(PdfTextExtractor.GetTextFromPage(reader, i));
                    }
                }

                // Check if the string exists and output results
                if (text.ToString().Contains(targetString))
                {
                    Console.WriteLine($"Success! Found the string: *{targetString}*");
                }
                else
                {
                    Console.WriteLine($"Sorry, the string *{targetString}* was not found in the PDF.");
                }
            }
            catch (Exception e)
            {
                Console.WriteLine("The file could not be read:");
                Console.WriteLine(e.Message);
            }
            Console.Read();
        }
    }
}

Option 2: Fine-Grained Matching (With Linq, As You Thought)

If you want to locate exactly where the string appears (including page numbers and position) and use the Linq syntax you mentioned, you'll need to use LocationTextExtractionStrategy to get individual text chunks instead of the full page text. This lets you filter and sort the chunks:

First, add the System.Linq namespace, then modify your code like this:

using iTextSharp.text.pdf;
using iTextSharp.text.pdf.parser;
using System;
using System.Collections.Generic;
using System.Linq;

namespace PDFReader
{
    internal class Class1
    {
        private static void Main(string[] args)
        {
            try
            {
                string targetString = "SpecifiedStringHere"; // Your target string here
                List<TextChunk> allTextChunks = new List<TextChunk>();

                using (PdfReader reader = new PdfReader(@"C:\Users\Me\Downloads\TestPDF.pdf"))
                {
                    for (int i = 1; i <= reader.NumberOfPages; i++)
                    {
                        // Use LocationTextExtractionStrategy to capture individual text chunks
                        var extractionStrategy = new LocationTextExtractionStrategy();
                        PdfTextExtractor.GetTextFromPage(reader, i, extractionStrategy);
                        allTextChunks.AddRange(extractionStrategy.GetLocationTextChunks());
                    }
                }

                // Filter, sort, and collect matching chunks (matches your Linq idea!)
                var matchingChunks = allTextChunks
                    .Where(chunk => chunk.Text.Contains(targetString))
                    .OrderBy(chunk => chunk.Location.Y) // PDF uses bottom-left origin, so Y increases upward
                    .Reverse() // Reverse to get top-to-bottom order
                    .ToList();

                // Output the results
                if (matchingChunks.Any())
                {
                    Console.WriteLine($"Found {matchingChunks.Count} occurrences of *{targetString}*:");
                    foreach (var chunk in matchingChunks)
                    {
                        Console.WriteLine($"Page {chunk.Location.PageNumber}: {chunk.Text}");
                    }
                }
                else
                {
                    Console.WriteLine($"No occurrences of *{targetString}* found in the PDF.");
                }
            }
            catch (Exception e)
            {
                Console.WriteLine("The file could not be read:");
                Console.WriteLine(e.Message);
            }
            Console.Read();
        }
    }
}

A Quick Note on Sorting:

PDFs use a coordinate system where the origin is the bottom-left corner of the page. That means higher Y-values correspond to positions higher up on the page. By using OrderBy(chunk => chunk.Location.Y).Reverse(), we're sorting the matching chunks from top to bottom of the document—just like you'd read it.

内容的提问来源于stack exchange,提问作者Wheelybob

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 20:28:00