Errno 22 Invalid argument路径报错,Python处理Excel URI批量复制PDF问题
问题原因
- 核心报错诱因是
shutil.copy()仅支持操作系统原生格式的文件路径,无法直接处理file://开头的URI格式路径,你读取到的业务表格URI没有经过转换直接传入,因此触发参数非法错误 - 你示例的
file://///XXX/2016/Report 140428 2.pdf是SMB网络共享路径的URI表示,Windows资源管理器内置了URI到原生UNC路径的自动转换逻辑,所以直接粘贴可以打开,Python标准库没有内置这个自动转换步骤
解决方案
你只需要在读取URI之后,先把URI转换为Windows可识别的原生路径即可,修正后的代码如下:
import pandas as pd import shutil import os from urllib.parse import unquote def uri_to_windows_path(uri): # 解码URI中可能存在的URL编码字符(比如空格被转义为%20的情况) decoded = unquote(uri) # 移除file://前缀 path = decoded.removeprefix('file://') # 替换斜杠为Windows路径标准的反斜杠 path = path.replace('/', '\\') # 处理UNC共享路径格式,保留开头的双反斜杠 return path.lstrip('\\') if not path.startswith('\\\\') else path xlsx = pd.ExcelFile(r"C:\Users\me\summary.xlsx") df = pd.read_excel(xlsx,"Sheet1",engine="openpyxl") clean = df.dropna(axis=1, how='all', inplace=False) dedupe = clean.drop_duplicates(["File Size","File Name"]) # 提前创建目标目录避免找不到路径报错 os.makedirs(r"C:\my new path", exist_ok=True) for index,row in dedupe.iterrows(): source_uri = row["URI"] source_path = uri_to_windows_path(source_uri) dest_dir = r"C:\my new path" # 拼接目标完整路径,避免同名文件覆盖 dest_path = os.path.join(dest_dir, os.path.basename(source_path)) # 增加异常捕获,单个文件报错不影响整体遍历 try: shutil.copy(source_path, dest_path) except Exception as e: print(f"处理文件{source_path}失败:{e}") print("pdfs written")
如果使用Python3.9以下版本,将
decoded.removeprefix('file://')替换为decoded.replace('file://', '', 1)即可兼容
内容的提问来源于stack exchange,提问作者Cragglecat
相关产品推荐
相关产品推荐

