You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何对Amazon S3 ls命令返回的文件按后缀分类统计

实现方案

首先你原有代码使用os.system仅能获取命令执行退出码,无法捕获输出内容,因此需要改用subprocess模块获取aws cli的执行结果,再做内容解析、后缀统计即可,具体实现如下:

import subprocess
import os
from collections import defaultdict

# 替换为目标存储桶地址
bucket_url = "s3://bucket"
cmd = f"aws s3 ls {bucket_url} --recursive --human-readable --summarize"

# 捕获命令输出
output = subprocess.check_output(cmd, shell=True, text=True)

ext_counter = defaultdict(int)
total_objects = 0
# 该变量可在后续代码中直接调用
total_size = ""

for line in output.splitlines():
    line = line.strip()
    if not line:
        continue
    # 解析汇总行
    if line.startswith("Total Objects:"):
        total_objects = int(line.split(":", 1)[1].strip())
    elif line.startswith("Total Size:"):
        total_size = line.split(":", 1)[1].strip()
    # 解析文件行统计后缀
    else:
        line_parts = line.split()
        if len(line_parts) < 4:
            continue
        file_name = line_parts[-1]
        # 获取带点的文件后缀
        file_ext = os.path.splitext(file_name)[1].lower()
        if file_ext:
            ext_counter[file_ext] += 1

# 按要求格式输出
print(f"Total Objects: {total_objects}")
print(f"Total Size: {total_size}")
for ext, count in sorted(ext_counter.items()):
    print(f"{ext} count: {count}")

注意事项

  • 运行前请确保本地已安装aws cli,且完成身份权限配置,可正常访问目标S3存储桶
  • 变量total_size存储了总大小值,可直接在后续代码中调用使用

内容的提问来源于stack exchange,提问作者Codu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 01:15:03