如何在命令行递归统计S3存储桶文件并按扩展名分组
解决方案
1. 按特定扩展名统计文件数量
直接过滤aws s3 ls的输出结果即可实现,比如统计.jpg格式文件:
aws s3 ls 's3://s3bucket_name/folder_name/' --recursive | grep "\.jpg$" | wc -l
--recursive:递归遍历存储桶目标路径下的所有文件grep "\.jpg$":精准匹配以.jpg结尾的文件名(转义.是避免它匹配任意字符)wc -l:统计匹配到的行数,对应文件总数
替换\.jpg$为其他扩展名就能统计对应格式,比如.png就用"\.png$"。
2. 一次性统计所有扩展名的文件数量
借助awk提取扩展名并分组统计:
aws s3 ls 's3://s3bucket_name/folder_name/' --recursive | awk ' /^[0-9]/ { split($4, ext, "."); if (length(ext) > 1) { ext_name = ext[length(ext)]; count[ext_name]++; } else { count["无扩展名"]++; } } END { for (e in count) { printf "%s: %d\n", e, count[e]; } } '
/^[0-9]/:只处理文件行(aws s3 ls输出中,文件行以日期数字开头,目录行不满足该规则)split($4, ext, "."):按.分割文件名,拆分出扩展名部分- 单独处理无扩展名的文件,归类到"无扩展名"分组
- 最后遍历统计结果,输出每个扩展名对应的文件数量
如果需要对结果排序,在命令末尾追加| sort即可:
aws s3 ls 's3://s3bucket_name/folder_name/' --recursive | awk ' /^[0-9]/ { split($4, ext, "."); if (length(ext) > 1) { ext_name = ext[length(ext)]; count[ext_name]++; } else { count["无扩展名"]++; } } END { for (e in count) { printf "%s: %d\n", e, count[e]; } } ' | sort
内容的提问来源于stack exchange,提问作者Jennifer Crosby
相关产品推荐
相关产品推荐

