You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Jupyter Lab环境下,如何编写代码读取PDF文件、提取指定标题下的表格并清理页码等冗余内容?

Hey David, great question! When dealing with structured PDFs that mix headings, subheadings, and embedded tables, there are a few robust Python-based approaches you can implement directly in Jupyter Lab. Let’s walk through the most practical ones, complete with code snippets to get you started:

方案1:PyPDF2 + Tabula (适合结构化程度较高的PDF)

This combo works well when your PDFs have clear, predictable headings and tables. PyPDF2 helps you locate which page your target heading lives on, and Tabula excels at extracting table data from specific pages or regions.

Step 1: Install required libraries

In your Jupyter Lab notebook, run this first:

!pip install pypdf2 tabula-py

Step 2: Locate the target heading’s page

First, we’ll scan the PDF to find which page contains your desired title:

from PyPDF2 import PdfReader

def find_heading_page(pdf_path, target_heading):
    reader = PdfReader(pdf_path)
    for page_num, page in enumerate(reader.pages, 1):
        text = page.extract_text()
        if target_heading in text:
            return page_num
    return None

# Example usage
pdf_path = "your_document.pdf"
target_title = "Quarterly Sales Data"
heading_page = find_heading_page(pdf_path, target_title)
print(f"Target heading found on page: {heading_page}")

Step 3: Extract and clean the table

Once you have the page number, use Tabula to pull the table, then filter out unwanted content like page numbers:

import tabula
import pandas as pd

# Extract table from the target page
tables = tabula.read_pdf(pdf_path, pages=heading_page, multiple_tables=False)
df = tables[0]

# Clean up page numbers or footer text (adjust the condition to match your PDF's format)
df = df[~df.apply(lambda row: any("Page" in str(cell) for cell in row), axis=1)]

# Reset index after filtering
df.reset_index(drop=True, inplace=True)

# View the cleaned table
display(df)

Pro tip: If the table spans multiple pages, adjust the pages parameter to include a range (e.g., pages=f"{heading_page}-{heading_page+2}").

方案2:Camelot (专为表格提取优化的工具)

Camelot is designed specifically for extracting tables from PDFs, and it gives you more control over filtering out noise like page numbers or stray text. It also outputs clean DataFrames by default.

Step 1: Install Camelot

!pip install camelot-py[cv]

Step 2: Extract and clean the table

import camelot

# Extract tables from the target page (use page range if needed)
tables = camelot.read_pdf(pdf_path, pages=str(heading_page))

# Get the first table (adjust index if multiple tables are present)
df = tables[0].df

# Remove rows containing page numbers (customize the pattern to match your PDF)
page_number_pattern = r"Page \d+"
df = df[~df.apply(lambda row: any(str(cell).match(page_number_pattern) for cell in row), axis=1)]

# Clean up any empty rows
df = df.dropna(how="all")

display(df)

Note: Camelot works best with PDFs that have clear table borders. If your tables are borderless, you might need to tweak the flavor parameter (try flavor="stream" instead of the default "lattice").

方案3:pdfplumber (精细控制文本和表格解析)

pdfplumber lets you parse PDFs at a granular level—you can locate the exact coordinates of your target heading, then extract only the table content below it. This is perfect for PDFs with messy layouts.

Step 1: Install pdfplumber

!pip install pdfplumber

Step 2: Locate heading coordinates and extract table

import pdfplumber

with pdfplumber.open(pdf_path) as pdf:
    page = pdf.pages[heading_page - 1]  # pdfplumber uses 0-indexed pages
    
    # Find the target heading's position
    words = page.extract_words()
    heading_coords = None
    for word in words:
        if word["text"] == target_title:
            heading_coords = (word["x0"], word["top"], word["x1"], word["bottom"])
            break
    
    if heading_coords:
        # Extract table below the heading (adjust the y-coordinate offset as needed)
        table = page.extract_table(
            table_settings={
                "vertical_strategy": "lines",
                "horizontal_strategy": "lines",
                "bbox": (0, heading_coords[3], page.width, page.height)  # Area below the heading
            }
        )
        
        # Convert to DataFrame
        df = pd.DataFrame(table[1:], columns=table[0])
        
        # Clean page numbers
        df = df[~df.apply(lambda row: any("Page" in str(cell) for cell in row), axis=1)]
        
        display(df)
    else:
        print("Target heading not found on the page.")

Pro tip: Use pdfplumber’s visual debugging tool (page.to_image().draw_rects(page.extract_words())) to see exactly where text is located, which helps adjust the bbox coordinates.

All these approaches work seamlessly in Jupyter Lab—you can tweak the filtering logic to match your PDF’s specific footer/page number format. Start with the simplest approach (PyPDF2 + Tabula) and move to Camelot or pdfplumber if you need more control.

内容的提问来源于stack exchange,提问作者David

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.27 15:27:41