You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用tabula-py读取Python课程表PDF时如何去除NaN值?工具替代建议及输出中‘...’含义解析

Troubleshooting tabula-py NaN Issues & Alternative Tools for PDF Timetable Parsing

Hey there! Let's tackle your problems step by step:

1. What do the ... mean in your output?

Those ellipses (...) are pandas' way of truncating long DataFrames when printing. Since your table has 8 columns, pandas hides the middle ones to keep the output readable. To see all columns, add this line before printing:

import pandas as pd
pd.set_option('display.max_columns', None)

This will force pandas to show every column in your DataFrame.

2. Fixing the excessive NaN values in tabula-py

Most NaNs here are likely caused by irregular table layouts (like merged cells or non-standard spacing) that tabula's default settings can't handle. Try these tweaks:

  • Use the right extraction mode:
    • If your timetable has visible borders, add lattice=True to read_pdf() — this tells tabula to parse the table based on grid lines:
      data = tabula.read_pdf(self.filename, pages='all', lattice=True)
      
    • If it's a borderless "stream" table, use stream=True instead, which parses based on text spacing.
  • Manual table area selection:
    Disable tabula's automatic table detection with guess=False, then define the exact area of your table (coordinates can be found using tabula's built-in GUI tool: run tabula.gui() in your terminal):
    data = tabula.read_pdf(self.filename, pages='all', guess=False, area=[20, 30, 700, 900])
    
  • Post-process with pandas:
    After extraction, clean up NaNs using pandas methods tailored to merged cells (common in timetables):
    df = data[0]  # Assuming your timetable is the first table
    # Fill merged cell values forward
    df = df.fillna(method='ffill')
    # Drop rows that are completely empty
    df = df.dropna(how='all')
    

3. Alternative tools to tabula-py

If tabula still isn't working well, here are reliable alternatives for parsing PDF timetables:

Camelot-py

Built specifically for PDF table extraction, Camelot offers better control over complex layouts and outputs an "accuracy score" for each table. It supports both lattice and stream modes:

import camelot
# Extract tables with lattice mode (for bordered tables)
tables = camelot.read_pdf('your_timetable.pdf', pages='all', flavor='lattice')
# Convert the first table to a DataFrame
df = tables[0].df

PDFplumber

A user-friendly tool built on PyMuPDF, PDFplumber excels at extracting structured data from PDFs, including handling merged cells. It gives you fine-grained control over text and table extraction:

import pdfplumber
import pandas as pd

with pdfplumber.open('your_timetable.pdf') as pdf:
    all_tables = []
    for page in pdf.pages:
        # Extract table with headers
        table = page.extract_table()
        if table:
            df = pd.DataFrame(table[1:], columns=table[0])
            all_tables.append(df)
# Combine tables from all pages
final_df = pd.concat(all_tables, ignore_index=True)

PyMuPDF (fitz)

For ultra-custom parsing, PyMuPDF lets you extract every text block with its exact coordinates. This is great if your timetable has a non-standard structure that other tools can't handle:

import fitz

doc = fitz.open('your_timetable.pdf')
for page in doc:
    # Extract all text blocks (each block has coordinates and text)
    blocks = page.get_text('blocks')
    for block in blocks:
        x0, y0, x1, y1, text = block[:5]
        # Process text based on its position (e.g., group by row/column)
        print(text.strip())

Desktop Tools (Adobe Acrobat Pro)

If you don't mind a manual step, Adobe Acrobat Pro can export PDF tables directly to Excel. Once exported, you can easily read the Excel file into pandas with pd.read_excel(). This often works surprisingly well for complex, non-standard tables.


内容的提问来源于stack exchange,提问作者Rishik

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 17:42:31