使用pypdfium2读取PDF后无法用shutil移动:文件未正常关闭问题
问题分析与解决方案
核心问题
代码中存在多处未正确释放的文件句柄,导致系统判定文件持续被占用:
pypdfium2.PdfDocument实例未被显式关闭,即便设置了autoclose=True,也可能因引用未被及时回收而保持文件锁定状态。pypdf.PdfReader长期持有原始PDF文件的句柄,加上Spyder环境的垃圾回收延迟,进一步加剧了文件占用问题。
修复步骤
1. 用with语句管理pypdfium2.PdfDocument生命周期
修改id_parsing方法,通过with语句自动释放文件句柄,替代依赖autoclose=True的被动回收:
def id_parsing(self, path, index): id = "" with pdfium.PdfDocument(path) as pdf: try: page = pdf[index].get_textpage() search_result = page.search("A285").get_next() if search_result is not None: x = search_result[0] idArr = page.get_text_bounded()[x+31:x+43].split() id = str(max(idArr, key=len)) except Exception: id = "-1" return id
注:原代码中通过字符串判断None的方式不严谨,直接判断对象是否为None更可靠。
2. 用with语句封装pypdf.PdfReader
在on_created方法中,将PdfReader放入with块,确保处理完成后立即释放原始文件句柄:
def on_created(self, event): print("Watchdog received created event - %s." % event.src_path) with PdfReader(event.src_path) as reader: lastId = "" id = "" startIndex = 0 for i in range(len(reader.pages) + 1): id = self.id_parsing(event.src_path, i) if id == "": id = lastId print("ID:" + id) if id != lastId and lastId != "": writer = PdfWriter() outputPdf = lastId + '.pdf' for page in range(startIndex, i): writer.add_page(reader.pages[page]) # 将文件写入操作移到循环外,避免重复锁定文件 with open(outputPdf, "wb") as f: writer.write(f) startIndex = i source = r'C:\Users\jkaplan\.spyder-py3' dest = r'C:\Users\jkaplan\Documents\Test\Processed Files' shutil.move(source + '\\' + outputPdf, dest) lastId = id dest2 = r'C:\Users\jkaplan\Documents\Test\Completed' # 此时原始文件已被释放,可安全移动 shutil.move(event.src_path, dest2)
注:原代码中writer.write(f)放在页面遍历循环内,会导致重复打开/写入文件,既降低效率又增加锁定风险,移到循环外一次性写入更合理。
3. 手动触发垃圾回收(可选)
若上述修改后仍存在锁定问题,可在shutil.move前手动调用垃圾回收,强制销毁未被引用的文件对象:
import gc # 在移动文件前添加 gc.collect()
关键注意事项
- 所有涉及文件操作的对象(PDF解析器、文件流)均需通过
with语句管理,确保资源自动释放。 - 避免在循环中重复打开同一文件,减少不必要的文件锁定。
- Spyder的交互式环境会保留对象引用,导致垃圾回收延迟,显式关闭资源比依赖自动回收更可靠。
内容的提问来源于stack exchange,提问作者Jacob
相关产品推荐
相关产品推荐

