如何使用iTextSharp读取无页眉页脚文件及移除现有PDF页眉页脚
Alright, let's tackle these two common iTextSharp tasks you're asking about—extracting text without headers/footers, and removing headers/footers from existing PDFs. I'll walk through each scenario with practical code examples and notes for edge cases (like files that might not have headers/footers at all).
1. Extract Text Excluding Headers and Footers
The default LocationTextExtractionStrategy pulls all text from a page, including headers and footers. To get only the main content, you'll need to create a custom extraction strategy that filters text based on its position on the page. Headers/footers are almost always in fixed top/bottom regions, so we can use coordinate bounds to exclude them.
First, define your custom strategy:
public class FilteredTextExtractionStrategy : LocationTextExtractionStrategy { private readonly float _headerHeight; private readonly float _footerHeight; private readonly float _pageHeight; public FilteredTextExtractionStrategy(float headerHeight, float footerHeight, float pageHeight) { _headerHeight = headerHeight; _footerHeight = footerHeight; _pageHeight = pageHeight; } public override void RenderText(TextRenderInfo renderInfo) { // Get the bottom y-coordinate of the text chunk float textBottom = renderInfo.GetDescentLine().GetStartPoint()[1]; // Only process text that's below the header and above the footer if (textBottom > _footerHeight && textBottom < (_pageHeight - _headerHeight)) { base.RenderText(renderInfo); } } }
Then use this strategy to extract text. Adjust the header/footer height values to match your actual PDF's layout:
using (PdfReader reader = new PdfReader("your-file-path.pdf")) { // Define header/footer heights (tweak these based on your document) float headerHeight = 50f; float footerHeight = 50f; int pageNumber = 1; // Target page number // Get the page's total height Rectangle pageSize = reader.GetPageSize(pageNumber); float pageHeight = pageSize.Height; // Extract text excluding header/footer regions string cleanText = PdfTextExtractor.GetTextFromPage( reader, pageNumber, new FilteredTextExtractionStrategy(headerHeight, footerHeight, pageHeight) ); // Use your clean text here }
Quick Note:
If the PDF has no headers/footers, this code still works flawlessly—it just extracts all the text since there's no content in the excluded regions. For variable PDFs, you could add logic to auto-detect header/footer bounds, but fixed values work for most standard documents.
2. Remove Headers and Footers from an Existing PDF
To permanently strip headers/footers, create a new PDF that copies the original page content but crops out the header/footer regions. We'll use PdfStamper for this straightforward approach:
using (PdfReader reader = new PdfReader("original-file.pdf")) using (PdfStamper stamper = new PdfStamper(reader, new FileStream("cleaned-file.pdf", FileMode.Create))) { float headerHeight = 50f; float footerHeight = 50f; for (int i = 1; i <= reader.NumberOfPages; i++) { Rectangle originalPageSize = reader.GetPageSize(i); // Calculate the new crop area: exclude top header and bottom footer Rectangle newCropBox = new Rectangle( originalPageSize.Left, originalPageSize.Bottom + footerHeight, originalPageSize.Right, originalPageSize.Top - headerHeight ); // Apply the crop box to the page stamper.GetUnderContent(i).SetCropBox(newCropBox); stamper.GetOverContent(i).SetCropBox(newCropBox); stamper.SetPageSize(newCropBox, i); } }
Alternative: Redraw Page Content (For Overlapping Content)
If cropping isn't enough (e.g., headers/footers overlap with main content), redraw the page content onto a new canvas, only including the middle region:
using (PdfReader reader = new PdfReader("original-file.pdf")) using (Document document = new Document()) using (PdfWriter writer = PdfWriter.GetInstance(document, new FileStream("cleaned-file.pdf", FileMode.Create))) { document.Open(); float headerHeight = 50f; float footerHeight = 50f; for (int i = 1; i <= reader.NumberOfPages; i++) { Rectangle originalPage = reader.GetPageSize(i); // Set new document size to exclude header/footer document.SetPageSize(new Rectangle( originalPage.Left, originalPage.Bottom + footerHeight, originalPage.Right, originalPage.Top - headerHeight )); document.NewPage(); // Import and redraw the original page content PdfImportedPage importedPage = writer.GetImportedPage(reader, i); // Offset content to skip footer region float yOffset = -footerHeight; writer.DirectContent.AddTemplate(importedPage, 0, yOffset); } document.Close(); }
Quick Note:
If the original PDF has no headers/footers, this code just creates an exact copy—so it's safe to run on all your target files regardless of their layout.
内容的提问来源于stack exchange,提问作者shmookh

