You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python基于PDF文本复制文件遇FileNotFoundError,文件存在却无法复制

解决PDF收据分类复制时的FileNotFoundError问题

问题场景

尝试根据PDF文本内容将收据文件分类复制到对应文件夹,程序能正确列出源文件夹中的文件,但执行复制操作时触发FileNotFoundError: [Errno 2] No such file or directory错误。

错误原因

核心问题是手动拼接路径时缺少路径分隔符:

  • 原代码中直接用ls_base + files拼接源文件路径,Windows系统下会生成类似D:\OneDrive\Documents\ReceiptsReceipt1.pdf的无效路径(源路径和文件名之间没有反斜杠分隔)
  • 目标路径ls_t1 + files、ls_t2 + files存在同样问题,导致系统无法识别目标路径

修复后的代码

import os
import re
import PyPDF2
import shutil

# 设置路径
ls_base = r"D:\OneDrive\Documents\Receipts" 
ls_t1 = r"D:\OneDrive\Documents\Receipts\Org1"
ls_t2 = r"D:\OneDrive\Documents\Receipts\Org2"

# 遍历源文件夹中的PDF文件
for filename in os.listdir(ls_base):
    if filename.endswith(".pdf"):
        # 用os.path.join正确拼接源文件路径
        source_path = os.path.join(ls_base, filename)
        
        # 读取PDF内容(用with语句自动管理文件对象)
        with open(source_path, 'rb') as pdfFileObj:
            pdfReader = PyPDF2.PdfReader(pdfFileObj)
            pageObj = pdfReader.pages[0]
            text = pageObj.extract_text()
            ls_name = re.findall(r'The Paint Shop', text)
        
        # 确定目标路径
        if ls_name:
            target_path = os.path.join(ls_t1, filename)
            print(ls_name, text.split('Items')[1].split('\n')[1], filename)
        else:
            target_path = os.path.join(ls_t2, filename)
        
        # 执行复制
        shutil.copy2(source_path, target_path)

关键修改说明

  1. 使用os.path.join()拼接路径:自动适配Windows的反斜杠和其他系统的路径分隔符,彻底避免手动拼接的分隔符错误
  2. 使用with语句管理文件对象:无需手动调用close(),代码更简洁且避免资源泄漏
  3. 移除不必要的os.chdir()操作:直接使用完整路径操作文件,避免目录切换带来的潜在路径混乱

内容的提问来源于stack exchange,提问作者Saleem

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 15:57:04