Python拆分PDF后删除临时目录报文件被占用问题解决方案
问题原因
你遇到的文件占用问题,核心是没有释放全所有PDF相关的文件句柄:你只尝试关闭了原始读入的PDF文件,但循环中通过Pdf.new()创建的所有拆分后新PDF对象,保存完成后没有执行关闭操作。这些对象会持有原始PDF的页面资源引用,Windows系统下不会自动回收这部分句柄,最终导致临时目录里的原PDF被锁定,无法删除。
你之前加的pdf.close()没有覆盖到所有需要关闭的对象,自然不生效;Pdf.merger.close()和你当前的拆分逻辑完全无关,没用是正常的。
修复方案
核心修复点
- 所有PDF对象(原始读入的PDF、拆分生成的每个新PDF),保存完成后必须显式调用
close()释放句柄 - 优先用
with上下文管理器处理PDF读写,哪怕代码运行中途报错,也会自动释放资源,不会残留文件锁 - 原代码最后保存最后一页PDF时存在路径拼接笔误、os.walk遍历存在变量名冲突问题,修复版里一并修正
修复后的核心代码
import os import shutil import time # 如果你用的是旧版PyPDF2库,把下面的导入替换为 from PyPDF2 import PdfReader, PdfWriter 即可,逻辑完全一致 from pypdf import Pdf run = 0 tempPath = r"E:\Intern\Programmieren\Python for Work\Testumgebung\temp" geoPath = r"E:\Intern\Programmieren\Python for Work\Testumgebung\Georeferenzieren" print("###Split up PDFs") file2pages = { 0: [0,1], 1: [1,2], 2: [2,3], 3: [3,4], 4: [4,5], } for root, directories, files in os.walk(tempPath): for file in files: run += 1 filePath = os.path.join(root, file) # 用上下文管理器打开原PDF,退出块时自动关闭释放句柄 with Pdf.open(filePath) as pdf: newPdfFiles = [Pdf.new() for _ in file2pages] newPdfIndex = 0 for n, page in enumerate(pdf.pages): if n in list(range(*file2pages[newPdfIndex])): newPdfFiles[newPdfIndex].pages.append(page) else: pdfPath = os.path.join(geoPath, str(newPdfIndex), file) outputFilename = f"{pdfPath}-{newPdfIndex}.pdf" newPdfFiles[newPdfIndex].save(outputFilename) # 保存完立刻关闭当前拆分出的PDF,释放资源 newPdfFiles[newPdfIndex].close() newPdfIndex += 1 newPdfFiles[newPdfIndex].pages.append(page) # 保存最后一个拆分出的PDF pdfPath = os.path.join(geoPath, str(newPdfIndex), file) outputFilename = f"{pdfPath}-{newPdfIndex}.pdf" newPdfFiles[newPdfIndex].save(outputFilename) # 关闭最后一个新PDF对象 newPdfFiles[newPdfIndex].close() print(f"Splitting up the {run}. PDF.") # 短暂等待系统完成句柄回收 time.sleep(0.5) print("\n#Removing temp folder...") shutil.rmtree(tempPath)
额外说明
不要靠加很长的sleep等待系统自动释放句柄,显式关闭所有打开的PDF对象才是根本解决方法,改完之后不会再出现文件被占用的报错。如果用旧版PyPDF2,对应把Pdf.open替换为PdfReader、Pdf.new替换为PdfWriter即可,关闭逻辑没有区别。
内容的提问来源于stack exchange,提问作者duxt3r__
相关产品推荐
相关产品推荐

