You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python拆分PDF后删除临时目录报文件被占用问题解决方案

问题原因

你遇到的文件占用问题,核心是没有释放全所有PDF相关的文件句柄:你只尝试关闭了原始读入的PDF文件,但循环中通过Pdf.new()创建的所有拆分后新PDF对象,保存完成后没有执行关闭操作。这些对象会持有原始PDF的页面资源引用,Windows系统下不会自动回收这部分句柄,最终导致临时目录里的原PDF被锁定,无法删除。
你之前加的pdf.close()没有覆盖到所有需要关闭的对象,自然不生效;Pdf.merger.close()和你当前的拆分逻辑完全无关,没用是正常的。

修复方案

核心修复点

  • 所有PDF对象(原始读入的PDF、拆分生成的每个新PDF),保存完成后必须显式调用close()释放句柄
  • 优先用with上下文管理器处理PDF读写,哪怕代码运行中途报错,也会自动释放资源,不会残留文件锁
  • 原代码最后保存最后一页PDF时存在路径拼接笔误、os.walk遍历存在变量名冲突问题,修复版里一并修正

修复后的核心代码

import os
import shutil
import time
# 如果你用的是旧版PyPDF2库,把下面的导入替换为 from PyPDF2 import PdfReader, PdfWriter 即可,逻辑完全一致
from pypdf import Pdf

run = 0
tempPath = r"E:\Intern\Programmieren\Python for Work\Testumgebung\temp"
geoPath = r"E:\Intern\Programmieren\Python for Work\Testumgebung\Georeferenzieren"

print("###Split up PDFs")

file2pages = {
    0: [0,1], 1: [1,2], 2: [2,3], 3: [3,4], 4: [4,5],
}

for root, directories, files in os.walk(tempPath):
    for file in files:
        run += 1
        filePath = os.path.join(root, file)
        # 用上下文管理器打开原PDF,退出块时自动关闭释放句柄
        with Pdf.open(filePath) as pdf:
            newPdfFiles = [Pdf.new() for _ in file2pages]
            newPdfIndex = 0

            for n, page in enumerate(pdf.pages):
                if n in list(range(*file2pages[newPdfIndex])):
                    newPdfFiles[newPdfIndex].pages.append(page)
                else:
                    pdfPath = os.path.join(geoPath, str(newPdfIndex), file)
                    outputFilename = f"{pdfPath}-{newPdfIndex}.pdf"
                    newPdfFiles[newPdfIndex].save(outputFilename)
                    # 保存完立刻关闭当前拆分出的PDF,释放资源
                    newPdfFiles[newPdfIndex].close()
                    newPdfIndex += 1
                    newPdfFiles[newPdfIndex].pages.append(page)

            # 保存最后一个拆分出的PDF
            pdfPath = os.path.join(geoPath, str(newPdfIndex), file)
            outputFilename = f"{pdfPath}-{newPdfIndex}.pdf"
            newPdfFiles[newPdfIndex].save(outputFilename)
            # 关闭最后一个新PDF对象
            newPdfFiles[newPdfIndex].close()
            print(f"Splitting up the {run}. PDF.")

# 短暂等待系统完成句柄回收
time.sleep(0.5)
print("\n#Removing temp folder...")
shutil.rmtree(tempPath)

额外说明

不要靠加很长的sleep等待系统自动释放句柄,显式关闭所有打开的PDF对象才是根本解决方法,改完之后不会再出现文件被占用的报错。如果用旧版PyPDF2,对应把Pdf.open替换为PdfReader、Pdf.new替换为PdfWriter即可,关闭逻辑没有区别。

内容的提问来源于stack exchange,提问作者duxt3r__

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 10:18:05