You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pypdfium2读取PDF后无法用shutil移动:文件未正常关闭问题

问题分析与解决方案

核心问题

代码中存在多处未正确释放的文件句柄,导致系统判定文件持续被占用:

  1. pypdfium2.PdfDocument 实例未被显式关闭,即便设置了 autoclose=True,也可能因引用未被及时回收而保持文件锁定状态。
  2. pypdf.PdfReader 长期持有原始PDF文件的句柄,加上Spyder环境的垃圾回收延迟,进一步加剧了文件占用问题。

修复步骤

1. 用with语句管理pypdfium2.PdfDocument生命周期

修改id_parsing方法,通过with语句自动释放文件句柄,替代依赖autoclose=True的被动回收:

def id_parsing(self, path, index):
    id = ""
    with pdfium.PdfDocument(path) as pdf:
        try:
            page = pdf[index].get_textpage()
            search_result = page.search("A285").get_next()
            if search_result is not None:
                x = search_result[0]
                idArr = page.get_text_bounded()[x+31:x+43].split()
                id = str(max(idArr, key=len))
        except Exception:
            id = "-1"
    return id

注:原代码中通过字符串判断None的方式不严谨,直接判断对象是否为None更可靠。

2. 用with语句封装pypdf.PdfReader

在on_created方法中,将PdfReader放入with块,确保处理完成后立即释放原始文件句柄:

def on_created(self, event):
    print("Watchdog received created event - %s." % event.src_path)
    with PdfReader(event.src_path) as reader:
        lastId = ""
        id = ""
        startIndex = 0

        for i in range(len(reader.pages) + 1):
            id = self.id_parsing(event.src_path, i)
            if id == "":
                id = lastId
            print("ID:" + id)
                
            if id != lastId and lastId != "":
                writer = PdfWriter()
                outputPdf = lastId + '.pdf'
                for page in range(startIndex, i):
                    writer.add_page(reader.pages[page])
                # 将文件写入操作移到循环外,避免重复锁定文件
                with open(outputPdf, "wb") as f:
                    writer.write(f)
                startIndex = i
                source = r'C:\Users\jkaplan\.spyder-py3'
                dest = r'C:\Users\jkaplan\Documents\Test\Processed Files'
                shutil.move(source + '\\' + outputPdf, dest)
                
            lastId = id
            
        dest2 = r'C:\Users\jkaplan\Documents\Test\Completed'
        # 此时原始文件已被释放,可安全移动
        shutil.move(event.src_path, dest2)

注:原代码中writer.write(f)放在页面遍历循环内,会导致重复打开/写入文件,既降低效率又增加锁定风险,移到循环外一次性写入更合理。

3. 手动触发垃圾回收(可选)

若上述修改后仍存在锁定问题,可在shutil.move前手动调用垃圾回收,强制销毁未被引用的文件对象:

import gc

# 在移动文件前添加
gc.collect()

关键注意事项

  • 所有涉及文件操作的对象(PDF解析器、文件流)均需通过with语句管理,确保资源自动释放。
  • 避免在循环中重复打开同一文件,减少不必要的文件锁定。
  • Spyder的交互式环境会保留对象引用,导致垃圾回收延迟,显式关闭资源比依赖自动回收更可靠。

内容的提问来源于stack exchange,提问作者Jacob

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.18 19:34:54