如何用Python读取多个含URL的JSON文件并下载对应资源
解决方案
1. 功能说明
在你已有的CSV转换逻辑基础上,新增Python实现的URL下载功能,直接嵌入Automator流程即可使用,无需额外安装依赖。
2. 完整Automator配置
保持原有配置不变:
- Shell =
/bin/bash - Pass input =
as arguments
替换原有Code为以下内容,原有CSV转换逻辑完全保留,新增了内嵌Python下载逻辑:
#!/bin/bash # 原有CSV转换逻辑 /usr/bin/perl -CSDA -w <<'EOF' - "$@" > ~/Desktop/out_"$(date '+%F_%H%M%S')".csv use strict; use JSON::Syck; $JSON::Syck::ImplicitUnicode = 1; # json node paths to extract my @paths = ('/upload_date', '/title', '/webpage_url'); for (@ARGV) { my $json; open(IN, "<", $_) or die "$!"; { local $/; $json = <IN>; } close IN; my $data = JSON::Syck::Load($json) or next; my @values = map { &json_node_at_path($data, $_) } @paths; { # output CSV spec # - field separator = SPACE # - record separator = LF # - every field is quoted local $, = qq( ); local $\ = qq(\n); print map { s/"/""/og; q(").$_.q("); } @values; } } sub json_node_at_path ($$) { # $ : (reference) json object # $ : (string) node path # # E.g. Given node path = '/abc/0/def', it returns either # $obj->{'abc'}->[0]->{'def'} if $obj->{'abc'} is ARRAY; or # $obj->{'abc'}->{'0'}->{'def'} if $obj->{'abc'} is HASH. my ($obj, $path) = @_; my $r = $obj; for ( map { /(^.+$)/ } split /\//, $path ) { if ( /^[0-9]+$/ && ref($r) eq 'ARRAY' ) { $r = $r->[$_]; } else { $r = $r->{$_}; } } return $r; } EOF # 新增Python下载逻辑 /usr/bin/python3 << 'EOF' -- "$@" import json import os import urllib.request import urllib.parse import sys from pathlib import Path # 下载文件保存目录,可自行修改 DOWNLOAD_PATH = os.path.expanduser("~/Desktop/下载的文件") Path(DOWNLOAD_PATH).mkdir(parents=True, exist_ok=True) for json_file in sys.argv[1:]: try: with open(json_file, 'r', encoding='utf-8') as f: json_data = json.load(f) target_url = json_data.get("webpage_url") file_title = json_data.get("title", "未命名文件") if not target_url: continue # 清洗文件名,去除非法字符 safe_title = "".join([c for c in file_title if c not in '/\\:*?"<>|']) # 自动识别文件后缀 url_suffix = os.path.splitext(urllib.parse.urlparse(target_url).path)[1] or ".html" save_full_path = os.path.join(DOWNLOAD_PATH, f"{safe_title}{url_suffix}") # 执行下载 urllib.request.urlretrieve(target_url, save_full_path) except Exception: # 单个文件出错不中断整个流程 continue EOF
3. 注意事项
- 下载的文件默认保存在桌面的
下载的文件文件夹中,可修改代码中DOWNLOAD_PATH变量自定义路径 - 脚本兼容10000个文件的批量处理,单个文件异常不会中断整体流程
- 所有依赖都是Mac系统自带组件,无需额外安装第三方包
内容的提问来源于stack exchange,提问作者Nothingtoseehere
相关产品推荐
相关产品推荐

