求PDF第78-79页表格文本提取方案(Windows 10环境)
Hey there! Let's work through your problem of extracting table content from pages 78-79 of a PDF and exporting it to a TXT file exactly how you need it. I'll focus on actively maintained tools that play nice with Windows 10, whether you prefer Python (PyCharm) or C# (Visual Studio 2017).
优先推荐:Python方案(PyCharm)
Python has some fantastic, well-maintained libraries for PDF processing that are easy to set up once you avoid common installation pitfalls.
1. PyMuPDF (fitz)
This library is super reliable, updated regularly, and handles both text and table extraction smoothly. It's also lightweight and works great on Windows.
Installation Fixes
If you hit issues with pip install pymupdf, try:
- Using PyCharm's built-in package manager: Go to Settings > Project > Python Interpreter, click the
+icon, search forpymupdf, and click Install Package (this avoids most network-related errors). - Or use a domestic PyPI mirror:
pip install -i https://pypi.tuna.tsinghua.edu.cn/simple pymupdf
Example Code
This code will extract tables from pages 78-79, and write each table row as a single line in your TXT file (matching your example format):
import fitz # PyMuPDF # Replace with your PDF path pdf_file = "your_document.pdf" output_txt = "extracted_table_content.txt" # Open the PDF with fitz.open(pdf_file) as doc: # PyMuPDF uses 0-indexed pages, so 78 = 77, 79 = 78 target_page_indices = [77, 78] with open(output_txt, "w", encoding="utf-8") as txt_file: for page_idx in target_page_indices: page = doc[page_idx] # Detect tables on the page tables = page.find_tables() for table in tables: # Loop through each row in the table for row in table.rows: # Clean up cell content and join with spaces cleaned_cells = [cell.strip() for cell in row if cell.strip()] row_text = " ".join(cleaned_cells) txt_file.write(row_text + "\n") print(f"Done! Check {output_txt} for your extracted content.")
If you actually want each cell on its own line (your wording was a bit conflicting with the example), just replace the row processing part with:
for cell in row: cleaned_cell = cell.strip() if cleaned_cell: txt_file.write(cleaned_cell + "\n")
2. pdfplumber
Another excellent choice, specifically built for precise text and table extraction. It's great for structured tables like the one you're working with.
Installation
Same as above: Use PyCharm's package manager or run:
pip install -i https://pypi.tuna.tsinghua.edu.cn/simple pdfplumber
Example Code
import pdfplumber pdf_file = "your_document.pdf" output_txt = "extracted_table_content.txt" with pdfplumber.open(pdf_file) as pdf: # pdfplumber uses 1-indexed pages directly for page_num in [78, 79]: page = pdf.pages[page_num - 1] # Extract all tables from the page tables = page.extract_tables() with open(output_txt, "a", encoding="utf-8") as txt_file: for table in tables: for row in table: # Convert cells to string, clean, and join cleaned_cells = [str(cell).strip() for cell in row if str(cell).strip()] row_text = " ".join(cleaned_cells) txt_file.write(row_text + "\n") print("Extraction completed successfully!")
C#方案(Visual Studio 2017)
If you prefer sticking with VS2017, PdfPig is an actively maintained .NET library that works well for PDF text/table extraction.
Installation
- Open your project in VS2017.
- Right-click the project > Manage NuGet Packages.
- Search for
PdfPigand click Install.
Example Code
This code extracts text from the target pages and formats it to match your example (adjust the row detection threshold if needed for your table):
using System; using System.IO; using UglyToad.PdfPig; using UglyToad.PdfPig.Content; namespace PdfTableExtractor { class Program { static void Main(string[] args) { string pdfPath = "your_document.pdf"; string outputPath = "extracted_table_content.txt"; int[] targetPages = { 78, 79 }; // 1-indexed pages using (var document = PdfDocument.Open(pdfPath)) { using (var writer = new StreamWriter(outputPath, false, System.Text.Encoding.UTF8)) { foreach (int pageNum in targetPages) { Page page = document.GetPage(pageNum); var words = page.GetWords(); string currentRow = ""; float lastYPosition = -1; // Group words into rows based on Y-coordinate (adjust threshold as needed) foreach (var word in words) { if (lastYPosition != -1 && Math.Abs(word.BoundingBox.Bottom - lastYPosition) > 2) { writer.WriteLine(currentRow.Trim()); currentRow = ""; } currentRow += word.Text + " "; lastYPosition = word.BoundingBox.Bottom; } // Write the last row of the page if (!string.IsNullOrWhiteSpace(currentRow)) { writer.WriteLine(currentRow.Trim()); } } } } Console.WriteLine("Extraction done! Check the output file."); } } }
Key Notes
- All these libraries are actively maintained, so you don't have to worry about outdated dependencies.
- If you still hit installation issues, double-check that your Python/.NET environment is up to date (for VS2017, make sure you have the latest NuGet package manager updates).
内容的提问来源于stack exchange,提问作者user14092875

