如何对Amazon S3 ls命令返回的文件按后缀分类统计
实现方案
首先你原有代码使用os.system仅能获取命令执行退出码,无法捕获输出内容,因此需要改用subprocess模块获取aws cli的执行结果,再做内容解析、后缀统计即可,具体实现如下:
import subprocess import os from collections import defaultdict # 替换为目标存储桶地址 bucket_url = "s3://bucket" cmd = f"aws s3 ls {bucket_url} --recursive --human-readable --summarize" # 捕获命令输出 output = subprocess.check_output(cmd, shell=True, text=True) ext_counter = defaultdict(int) total_objects = 0 # 该变量可在后续代码中直接调用 total_size = "" for line in output.splitlines(): line = line.strip() if not line: continue # 解析汇总行 if line.startswith("Total Objects:"): total_objects = int(line.split(":", 1)[1].strip()) elif line.startswith("Total Size:"): total_size = line.split(":", 1)[1].strip() # 解析文件行统计后缀 else: line_parts = line.split() if len(line_parts) < 4: continue file_name = line_parts[-1] # 获取带点的文件后缀 file_ext = os.path.splitext(file_name)[1].lower() if file_ext: ext_counter[file_ext] += 1 # 按要求格式输出 print(f"Total Objects: {total_objects}") print(f"Total Size: {total_size}") for ext, count in sorted(ext_counter.items()): print(f"{ext} count: {count}")
注意事项
- 运行前请确保本地已安装aws cli,且完成身份权限配置,可正常访问目标S3存储桶
- 变量
total_size存储了总大小值,可直接在后续代码中调用使用
内容的提问来源于stack exchange,提问作者Codu
相关产品推荐
相关产品推荐

