You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python从多个JSON文件的URL批量下载PDF及现有代码问题修复

现有代码问题汇总
  • 基础语法错误:Python导入关键字为小写import,你写的Import会直接触发语法报错
  • 混用命令行工具与Python代码:in2csv是csvkit提供的命令行工具,不能直接写在Python脚本中执行
  • 文件读取逻辑错误:open('*.json', 'r')不支持通配符匹配,无法一次性读取目录下所有JSON文件
  • 数据结构认知错误:你的JSON示例是顶层对象结构,不是列表,urls_dict = urls_dict[0]会直接抛出索引错误
  • 缺失核心逻辑:没有PDF下载请求、流写入的相关代码,f.write(r.pdf)属于无效代码,变量r未定义
  • 无异常处理逻辑:50万量级的文件处理极易遇到JSON损坏、下载链接过期、网络波动等问题,无异常捕获会导致程序中途崩溃,进度全部丢失
修复实现方案

你不需要额外转CSV,直接遍历所有JSON文件提取链接下载即可,效率更高,下面是可直接运行的代码:

import json
import os
import requests
from pathlib import Path
from concurrent.futures import ThreadPoolExecutor, as_completed

# 配置项
JSON_DIR = Path("/Users/MyComputer/Desktop/self_mailers")
PDF_SAVE_DIR = Path("/Users/MyComputer/Desktop/downloaded_pdfs")
MAX_WORKERS = 10  # 可根据你的网络情况调整并发数
TIMEOUT = 30

# 创建保存目录
PDF_SAVE_DIR.mkdir(exist_ok=True)

def download_pdf(json_path):
    try:
        # 读取JSON文件
        with open(json_path, 'r', encoding='utf-8') as f:
            data = json.load(f)
        # 提取PDF链接和唯一ID(用ID当文件名避免重复)
        pdf_url = data.get('press_proof')
        pdf_id = data.get('id')
        if not pdf_url or not pdf_id:
            return f"跳过文件{json_path.name}:缺失必要字段"
        save_path = PDF_SAVE_DIR / f"{pdf_id}.pdf"
        if save_path.exists():
            return f"跳过文件{pdf_id}:已下载"
        # 下载PDF
        resp = requests.get(pdf_url, timeout=TIMEOUT, stream=True)
        resp.raise_for_status()
        with open(save_path, 'wb') as f:
            for chunk in resp.iter_content(chunk_size=8192):
                f.write(chunk)
        return f"下载完成:{pdf_id}"
    except Exception as e:
        return f"处理文件{json_path.name}失败:{str(e)}"

if __name__ == "__main__":
    # 遍历所有JSON文件
    json_files = list(JSON_DIR.glob("*.json"))
    total = len(json_files)
    print(f"共找到{total}个JSON文件,开始处理")
    # 多线程下载
    success = 0
    fail = 0
    with ThreadPoolExecutor(max_workers=MAX_WORKERS) as executor:
        futures = [executor.submit(download_pdf, jp) for jp in json_files]
        for future in as_completed(futures):
            res = future.result()
            print(res)
            if "下载完成" in res or ("跳过" in res and "已下载" in res):
                success +=1
            else:
                fail +=1
    print(f"全部处理完成,成功:{success},失败:{fail}")

如果你确实需要先转CSV再处理,可以在终端批量执行in2csv命令批量转换所有JSON,再修改上面的代码读取CSV的press_proof和id字段即可。

内容的提问来源于stack exchange,提问作者Nothingtoseehere

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 16:57:02