使用PyMuPDF提取PDF表格时函数调用致Jupyter Kernel崩溃
PyMuPDF封装函数后Jupyter内核崩溃问题
问题描述
单独运行PyMuPDF提取PDF表格的脚本时完全正常,能成功获取目标表格,但将逻辑封装成函数调用时,Jupyter内核直接崩溃,单元格无限运行,日志显示:
11:46:41.424 [error] Disposing session as kernel process died ExitCode: 3221225477, Reason:
已尝试打印调试、更新PyMuPDF,且PDF仅两页,排除内存问题。
正常运行的脚本
import pandas as pd import pymupdf doc = pymupdf.open(filePath) pageNum = 1 page = doc[pageNum] # "afuhfahuashsadhhadasdhukads" 是故意设置的不存在文本,让left_anchor为空列表 left_anchor = page.search_for("afuhfahuashsadhhadasdhukads") top_anchor = page.search_for("Characteristics") right_anchor = page.search_for("Market Allocation") bottom_anchor = page.search_for("Top 10 holdings") x0 = left_anchor[0].x1 if len(left_anchor) > 0 else page.rect.x0 y0 = top_anchor[0].y1 if len(top_anchor) > 0 else page.rect.y0 x1 = right_anchor[0].x0 if len(right_anchor) > 0 else page.rect.x1 y1 = bottom_anchor[0].y0 if len(bottom_anchor) > 0 else page.rect.y1 clip = pymupdf.Rect(x0, y0, x1, y1) find_table = page.find_tables( clip = clip, vertical_strategy = "text", horizontal_strategy = "lines" ) table = find_table.tables[0].to_pandas() table
导致崩溃的函数及调用
import pymupdf def get_table(doc, page_num, left_anchor_text=None, top_anchor_text=None, right_anchor_text=None, bottom_anchor_text=None, vertical_strategy=None, horizontal_strategy=None): page = doc[page_num] left_anchor = page.search_for(left_anchor_text) top_anchor = page.search_for(top_anchor_text) right_anchor = page.search_for(right_anchor_text) bottom_anchor = page.search_for(bottom_anchor_text) x0 = left_anchor[0].x1 if len(left_anchor) > 0 else page.rect.x0 y0 = top_anchor[0].y1 if len(top_anchor) > 0 else page.rect.y0 x1 = right_anchor[0].x0 if len(right_anchor) > 0 else page.rect.x1 y1 = bottom_anchor[0].y0 if len(bottom_anchor) > 0 else page.rect.y1 clip = pymupdf.Rect(x0, y0, x1, y1) ft = page.find_tables( clip = clip, vertical_strategy = vertical_strategy if vertical_strategy else "lines", horizontal_strategy = horizontal_strategy if horizontal_strategy else "lines", ) table = ft.tables[0].to_pandas() return table document = pymupdf.open(filePath) get_table( doc=document, page_num=1, top_anchor_text="Performance", bottom_anchor_text="Calendar Year Performance", vertical_strategy="text", horizontal_strategy="lines" )
问题分析
ExitCode 3221225477对应Windows系统的访问违规错误(Access Violation),通常是代码试图访问未授权的内存地址,比如空指针引用。对比两段代码,核心差异在于:
- 函数版本中,
left_anchor_text、right_anchor_text默认传None(调用时未传入这两个参数),而单独脚本中是传了一个不存在的字符串 - PyMuPDF的
page.search_for方法传入None时,可能触发底层C扩展的空指针错误,直接导致进程崩溃
解决方案
修改函数,确保传给page.search_for的参数永远是字符串类型,未传入的锚点文本用空字符串替代None,同时增加边界检查避免索引越界:
import pymupdf import pandas as pd def get_table(doc, page_num, left_anchor_text="", top_anchor_text="", right_anchor_text="", bottom_anchor_text="", vertical_strategy="lines", horizontal_strategy="lines"): page = doc[page_num] # 处理空锚点文本,避免传入None触发底层错误 left_anchor = page.search_for(left_anchor_text) if left_anchor_text else [] top_anchor = page.search_for(top_anchor_text) if top_anchor_text else [] right_anchor = page.search_for(right_anchor_text) if right_anchor_text else [] bottom_anchor = page.search_for(bottom_anchor_text) if bottom_anchor_text else [] x0 = left_anchor[0].x1 if len(left_anchor) > 0 else page.rect.x0 y0 = top_anchor[0].y1 if len(top_anchor) > 0 else page.rect.y0 x1 = right_anchor[0].x0 if len(right_anchor) > 0 else page.rect.x1 y1 = bottom_anchor[0].y0 if len(bottom_anchor) > 0 else page.rect.y1 clip = pymupdf.Rect(x0, y0, x1, y1) ft = page.find_tables( clip = clip, vertical_strategy = vertical_strategy, horizontal_strategy = horizontal_strategy, ) # 检查是否找到表格,避免索引越界 if ft.tables: return ft.tables[0].to_pandas() else: return pd.DataFrame() # 未找到表格返回空DataFrame document = pymupdf.open(filePath) get_table( doc=document, page_num=1, top_anchor_text="Performance", bottom_anchor_text="Calendar Year Performance", vertical_strategy="text", horizontal_strategy="lines" )
额外优化说明
- 函数参数默认值直接设为合理的默认策略,简化代码逻辑
- 明确处理空锚点文本的情况,直接返回空列表,避免调用
search_for(None) - 增加表格存在性检查,防止
ft.tables[0]索引越界报错
内容的提问来源于stack exchange,提问作者OBramalam
相关产品推荐
相关产品推荐

