You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Errno 22 Invalid argument路径报错,Python处理Excel URI批量复制PDF问题

问题原因
  • 核心报错诱因是shutil.copy()仅支持操作系统原生格式的文件路径,无法直接处理file://开头的URI格式路径,你读取到的业务表格URI没有经过转换直接传入,因此触发参数非法错误
  • 你示例的file://///XXX/2016/Report 140428 2.pdf是SMB网络共享路径的URI表示,Windows资源管理器内置了URI到原生UNC路径的自动转换逻辑,所以直接粘贴可以打开,Python标准库没有内置这个自动转换步骤
解决方案

你只需要在读取URI之后,先把URI转换为Windows可识别的原生路径即可,修正后的代码如下:

import pandas as pd
import shutil
import os
from urllib.parse import unquote

def uri_to_windows_path(uri):
    # 解码URI中可能存在的URL编码字符(比如空格被转义为%20的情况)
    decoded = unquote(uri)
    # 移除file://前缀
    path = decoded.removeprefix('file://')
    # 替换斜杠为Windows路径标准的反斜杠
    path = path.replace('/', '\\')
    # 处理UNC共享路径格式,保留开头的双反斜杠
    return path.lstrip('\\') if not path.startswith('\\\\') else path

xlsx = pd.ExcelFile(r"C:\Users\me\summary.xlsx")
df = pd.read_excel(xlsx,"Sheet1",engine="openpyxl")
clean = df.dropna(axis=1, how='all', inplace=False)
dedupe = clean.drop_duplicates(["File Size","File Name"])
# 提前创建目标目录避免找不到路径报错
os.makedirs(r"C:\my new path", exist_ok=True)

for index,row in dedupe.iterrows():
    source_uri = row["URI"]
    source_path = uri_to_windows_path(source_uri)
    dest_dir = r"C:\my new path"
    # 拼接目标完整路径,避免同名文件覆盖
    dest_path = os.path.join(dest_dir, os.path.basename(source_path))
    # 增加异常捕获,单个文件报错不影响整体遍历
    try:
        shutil.copy(source_path, dest_path)
    except Exception as e:
        print(f"处理文件{source_path}失败:{e}")

print("pdfs written")

如果使用Python3.9以下版本,将decoded.removeprefix('file://')替换为decoded.replace('file://', '', 1)即可兼容

内容的提问来源于stack exchange,提问作者Cragglecat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 03:15:03