使用tabula-py读取Python课程表PDF时如何去除NaN值?工具替代建议及输出中‘...’含义解析
Hey there! Let's tackle your problems step by step:
1. What do the ... mean in your output?
Those ellipses (...) are pandas' way of truncating long DataFrames when printing. Since your table has 8 columns, pandas hides the middle ones to keep the output readable. To see all columns, add this line before printing:
import pandas as pd pd.set_option('display.max_columns', None)
This will force pandas to show every column in your DataFrame.
2. Fixing the excessive NaN values in tabula-py
Most NaNs here are likely caused by irregular table layouts (like merged cells or non-standard spacing) that tabula's default settings can't handle. Try these tweaks:
- Use the right extraction mode:
- If your timetable has visible borders, add
lattice=Truetoread_pdf()— this tells tabula to parse the table based on grid lines:data = tabula.read_pdf(self.filename, pages='all', lattice=True) - If it's a borderless "stream" table, use
stream=Trueinstead, which parses based on text spacing.
- If your timetable has visible borders, add
- Manual table area selection:
Disable tabula's automatic table detection withguess=False, then define the exact area of your table (coordinates can be found using tabula's built-in GUI tool: runtabula.gui()in your terminal):data = tabula.read_pdf(self.filename, pages='all', guess=False, area=[20, 30, 700, 900]) - Post-process with pandas:
After extraction, clean up NaNs using pandas methods tailored to merged cells (common in timetables):df = data[0] # Assuming your timetable is the first table # Fill merged cell values forward df = df.fillna(method='ffill') # Drop rows that are completely empty df = df.dropna(how='all')
3. Alternative tools to tabula-py
If tabula still isn't working well, here are reliable alternatives for parsing PDF timetables:
Camelot-py
Built specifically for PDF table extraction, Camelot offers better control over complex layouts and outputs an "accuracy score" for each table. It supports both lattice and stream modes:
import camelot # Extract tables with lattice mode (for bordered tables) tables = camelot.read_pdf('your_timetable.pdf', pages='all', flavor='lattice') # Convert the first table to a DataFrame df = tables[0].df
PDFplumber
A user-friendly tool built on PyMuPDF, PDFplumber excels at extracting structured data from PDFs, including handling merged cells. It gives you fine-grained control over text and table extraction:
import pdfplumber import pandas as pd with pdfplumber.open('your_timetable.pdf') as pdf: all_tables = [] for page in pdf.pages: # Extract table with headers table = page.extract_table() if table: df = pd.DataFrame(table[1:], columns=table[0]) all_tables.append(df) # Combine tables from all pages final_df = pd.concat(all_tables, ignore_index=True)
PyMuPDF (fitz)
For ultra-custom parsing, PyMuPDF lets you extract every text block with its exact coordinates. This is great if your timetable has a non-standard structure that other tools can't handle:
import fitz doc = fitz.open('your_timetable.pdf') for page in doc: # Extract all text blocks (each block has coordinates and text) blocks = page.get_text('blocks') for block in blocks: x0, y0, x1, y1, text = block[:5] # Process text based on its position (e.g., group by row/column) print(text.strip())
Desktop Tools (Adobe Acrobat Pro)
If you don't mind a manual step, Adobe Acrobat Pro can export PDF tables directly to Excel. Once exported, you can easily read the Excel file into pandas with pd.read_excel(). This often works surprisingly well for complex, non-standard tables.
内容的提问来源于stack exchange,提问作者Rishik

