You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python读取多个含URL的JSON文件并下载对应资源

解决方案

1. 功能说明

在你已有的CSV转换逻辑基础上,新增Python实现的URL下载功能,直接嵌入Automator流程即可使用,无需额外安装依赖。

2. 完整Automator配置

保持原有配置不变:

  • Shell = /bin/bash
  • Pass input = as arguments

替换原有Code为以下内容,原有CSV转换逻辑完全保留,新增了内嵌Python下载逻辑:

#!/bin/bash

# 原有CSV转换逻辑
/usr/bin/perl -CSDA -w <<'EOF' - "$@" > ~/Desktop/out_"$(date '+%F_%H%M%S')".csv
use strict;
use JSON::Syck;
$JSON::Syck::ImplicitUnicode = 1;

# json node paths to extract
 my @paths = ('/upload_date', '/title', '/webpage_url');

for (@ARGV) {
    my $json;
    open(IN, "<", $_) or die "$!";
    {
        local $/; 
        $json = <IN>;
    }
    close IN;
    my $data = JSON::Syck::Load($json) or next;
    my @values = map { &json_node_at_path($data, $_) } @paths;
    {
        #   output CSV spec
        #   - field separator = SPACE
        #   - record separator = LF
        #   - every field is quoted
        local $, = qq( );
        local $\ = qq(\n);
        print map { s/"/""/og; q(").$_.q("); } @values;
    }
}

sub json_node_at_path ($$) {
    #   $ : (reference) json object
    #   $ : (string) node path
    # 
    #   E.g. Given node path = '/abc/0/def', it returns either
    #       $obj->{'abc'}->[0]->{'def'}   if $obj->{'abc'} is ARRAY; or
    #       $obj->{'abc'}->{'0'}->{'def'} if $obj->{'abc'} is HASH.
    my ($obj, $path) = @_;  
    my $r = $obj;
    for ( map { /(^.+$)/ } split /\//, $path ) {
        if ( /^[0-9]+$/ && ref($r) eq 'ARRAY' ) {
        $r = $r->[$_];
        }
        else {
             $r = $r->{$_};
        }
    }
    return $r;
}
EOF

# 新增Python下载逻辑
/usr/bin/python3 << 'EOF' -- "$@"
import json
import os
import urllib.request
import urllib.parse
import sys
from pathlib import Path

# 下载文件保存目录,可自行修改
DOWNLOAD_PATH = os.path.expanduser("~/Desktop/下载的文件")
Path(DOWNLOAD_PATH).mkdir(parents=True, exist_ok=True)

for json_file in sys.argv[1:]:
    try:
        with open(json_file, 'r', encoding='utf-8') as f:
            json_data = json.load(f)
        target_url = json_data.get("webpage_url")
        file_title = json_data.get("title", "未命名文件")
        if not target_url:
            continue
        # 清洗文件名,去除非法字符
        safe_title = "".join([c for c in file_title if c not in '/\\:*?"<>|'])
        # 自动识别文件后缀
        url_suffix = os.path.splitext(urllib.parse.urlparse(target_url).path)[1] or ".html"
        save_full_path = os.path.join(DOWNLOAD_PATH, f"{safe_title}{url_suffix}")
        # 执行下载
        urllib.request.urlretrieve(target_url, save_full_path)
    except Exception:
        # 单个文件出错不中断整个流程
        continue
EOF

3. 注意事项

  • 下载的文件默认保存在桌面的下载的文件文件夹中,可修改代码中DOWNLOAD_PATH变量自定义路径
  • 脚本兼容10000个文件的批量处理,单个文件异常不会中断整体流程
  • 所有依赖都是Mac系统自带组件,无需额外安装第三方包

内容的提问来源于stack exchange,提问作者Nothingtoseehere

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 21:00:05