You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求PDF第78-79页表格文本提取方案(Windows 10环境)

解决PDF表格提取并导出为TXT的方案

Hey there! Let's work through your problem of extracting table content from pages 78-79 of a PDF and exporting it to a TXT file exactly how you need it. I'll focus on actively maintained tools that play nice with Windows 10, whether you prefer Python (PyCharm) or C# (Visual Studio 2017).

优先推荐:Python方案(PyCharm)

Python has some fantastic, well-maintained libraries for PDF processing that are easy to set up once you avoid common installation pitfalls.

1. PyMuPDF (fitz)

This library is super reliable, updated regularly, and handles both text and table extraction smoothly. It's also lightweight and works great on Windows.

Installation Fixes

If you hit issues with pip install pymupdf, try:

  • Using PyCharm's built-in package manager: Go to Settings > Project > Python Interpreter, click the + icon, search for pymupdf, and click Install Package (this avoids most network-related errors).
  • Or use a domestic PyPI mirror:
    pip install -i https://pypi.tuna.tsinghua.edu.cn/simple pymupdf
    

Example Code

This code will extract tables from pages 78-79, and write each table row as a single line in your TXT file (matching your example format):

import fitz  # PyMuPDF

# Replace with your PDF path
pdf_file = "your_document.pdf"
output_txt = "extracted_table_content.txt"

# Open the PDF
with fitz.open(pdf_file) as doc:
    # PyMuPDF uses 0-indexed pages, so 78 = 77, 79 = 78
    target_page_indices = [77, 78]
    
    with open(output_txt, "w", encoding="utf-8") as txt_file:
        for page_idx in target_page_indices:
            page = doc[page_idx]
            # Detect tables on the page
            tables = page.find_tables()
            
            for table in tables:
                # Loop through each row in the table
                for row in table.rows:
                    # Clean up cell content and join with spaces
                    cleaned_cells = [cell.strip() for cell in row if cell.strip()]
                    row_text = " ".join(cleaned_cells)
                    txt_file.write(row_text + "\n")

print(f"Done! Check {output_txt} for your extracted content.")

If you actually want each cell on its own line (your wording was a bit conflicting with the example), just replace the row processing part with:

for cell in row:
    cleaned_cell = cell.strip()
    if cleaned_cell:
        txt_file.write(cleaned_cell + "\n")

2. pdfplumber

Another excellent choice, specifically built for precise text and table extraction. It's great for structured tables like the one you're working with.

Installation

Same as above: Use PyCharm's package manager or run:

pip install -i https://pypi.tuna.tsinghua.edu.cn/simple pdfplumber

Example Code

import pdfplumber

pdf_file = "your_document.pdf"
output_txt = "extracted_table_content.txt"

with pdfplumber.open(pdf_file) as pdf:
    # pdfplumber uses 1-indexed pages directly
    for page_num in [78, 79]:
        page = pdf.pages[page_num - 1]
        # Extract all tables from the page
        tables = page.extract_tables()
        
        with open(output_txt, "a", encoding="utf-8") as txt_file:
            for table in tables:
                for row in table:
                    # Convert cells to string, clean, and join
                    cleaned_cells = [str(cell).strip() for cell in row if str(cell).strip()]
                    row_text = " ".join(cleaned_cells)
                    txt_file.write(row_text + "\n")

print("Extraction completed successfully!")

C#方案(Visual Studio 2017)

If you prefer sticking with VS2017, PdfPig is an actively maintained .NET library that works well for PDF text/table extraction.

Installation

  1. Open your project in VS2017.
  2. Right-click the project > Manage NuGet Packages.
  3. Search for PdfPig and click Install.

Example Code

This code extracts text from the target pages and formats it to match your example (adjust the row detection threshold if needed for your table):

using System;
using System.IO;
using UglyToad.PdfPig;
using UglyToad.PdfPig.Content;

namespace PdfTableExtractor
{
    class Program
    {
        static void Main(string[] args)
        {
            string pdfPath = "your_document.pdf";
            string outputPath = "extracted_table_content.txt";
            int[] targetPages = { 78, 79 }; // 1-indexed pages

            using (var document = PdfDocument.Open(pdfPath))
            {
                using (var writer = new StreamWriter(outputPath, false, System.Text.Encoding.UTF8))
                {
                    foreach (int pageNum in targetPages)
                    {
                        Page page = document.GetPage(pageNum);
                        var words = page.GetWords();
                        string currentRow = "";
                        float lastYPosition = -1;

                        // Group words into rows based on Y-coordinate (adjust threshold as needed)
                        foreach (var word in words)
                        {
                            if (lastYPosition != -1 && Math.Abs(word.BoundingBox.Bottom - lastYPosition) > 2)
                            {
                                writer.WriteLine(currentRow.Trim());
                                currentRow = "";
                            }
                            currentRow += word.Text + " ";
                            lastYPosition = word.BoundingBox.Bottom;
                        }

                        // Write the last row of the page
                        if (!string.IsNullOrWhiteSpace(currentRow))
                        {
                            writer.WriteLine(currentRow.Trim());
                        }
                    }
                }
            }

            Console.WriteLine("Extraction done! Check the output file.");
        }
    }
}

Key Notes

  • All these libraries are actively maintained, so you don't have to worry about outdated dependencies.
  • If you still hit installation issues, double-check that your Python/.NET environment is up to date (for VS2017, make sure you have the latest NuGet package manager updates).

内容的提问来源于stack exchange,提问作者user14092875

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 00:07:39