You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按项目统计指定子目录下的WAV文件数量?

问题需求
  • 按项目目录统计其下Files for analysis子目录中的所有WAV文件总数,忽略Backup目录内的文件
  • 项目目录结构:每个项目根目录下包含Files for analysis和Backup两个子目录,WAV文件存储在Files for analysis下的Round/Box层级子文件夹中
  • 示例文件路径:S:/sound_files/2024/R/testfolder/Project 1/Files for analysis/Round 1/Box 1/1.WAV

文件夹结构示例

| Box 1 --- 1.WAV, 2.WAV, 3.WAV
                                  | Round 1 --- | Box 2 --- 1.WAV, 2.WAV, 3.WAV
            | Files for analysis -              
            |                     | Round 2 --- | Box 3 --- 1.WAV, 2.WAV, 3.WAV
            |                                   | Box 4 --- 1.WAV, 2.WAV, 3.WAV
Project 1 --
            |                                   | Box 1 --- 1.WAV, 2.WAV, 3.WAV
            |                     | Round 1 --- | Box 2 --- 1.WAV, 2.WAV, 3.WAV         
            | Backup  ------------                     
                                  | Round 2 --- | Box 3 --- 1.WAV, 2.WAV, 3.WAV
                                                | Box 4 --- 1.WAV, 2.WAV, 3.WAV

                                                | Box 5 --- 1.WAV, 2.WAV, 3.WAV
                                  | Round 1 --- | Box 6 --- 1.WAV, 2.WAV, 3.WAV
            | Files for analysis -              
            |                     | Round 2 --- | Box 7 --- 1.WAV, 2.WAV, 3.WAV
            |                                   | Box 8 --- 1.WAV, 2.WAV, 3.WAV
Project 2 --
            |                                   | Box 5 --- 1.WAV, 2.WAV, 3.WAV
            |                     | Round 1 --- | Box 6 --- 1.WAV, 2.WAV, 3.WAV         
            | Backup  ------------                     
                                  | Round 2 --- | Box 7 --- 1.WAV, 2.WAV, 3.WAV
                                                | Box 8 --- 1.WAV, 2.WAV, 3.WAV

现有脚本及问题

现有R脚本

main <- "S:/sound_files/2024/R/testfolder"

## 列出所有文件夹 
dirs <- list.dirs(main, full.names = TRUE, recursive=TRUE)

## 列出顶级项目文件夹
only_mains <- dirs[lengths(strsplit(dirs, "/")) == 6 ] 

## 获取包含"Files for analysis"的文件夹
dir_files_for_analysis <- dirs[lengths(strsplit(dirs, "/")) == 7 ]
dir_files_for_analysis <- grep("Files for analysis", dir_files_for_analysis, value = TRUE) 

## 列出"Files for analysis"下的所有WAV文件
files <- list.files(dir_files_for_analysis, pattern = ".WAV", recursive = TRUE, full.names = TRUE) 

length(files) ## 统计WAV文件总数

## 按文件所在子目录分组
dir_list <- split(files, dirname(files)) 

files_in_folder <- sapply(dir_list, length)

head(files_in_folder)

当前输出(Box级统计)

S:/sound_files/2024/R/testfolder/Project 1/Files for analysis/Round 1/Box 1                                                                                                 
20

S:/sound_files/2024/R/testfolder/Project 1/Files for analysis/Round 2/Box 2                                                                                                              
19

S:/sound_files/2024/R/testfolder/Project 2/Files for analysis/Round 1/Box 3                                                                                                            
20

S:/sound_files/2024/R/testfolder/Project 2/Files for analysis/Round 2/Box 4                                                                                                            
20

期望输出(项目级统计)

S:/sound_files/2024/R/testfolder/Project 1   39 files
S:/sound_files/2024/R/testfolder/Project 2   40 files

解决方案(修改后的R脚本)

main <- "S:/sound_files/2024/R/testfolder"

# 递归获取所有WAV文件,同时过滤仅保留Files for analysis目录下的内容
files <- list.files(
  path = main,
  pattern = "\\.WAV$",  # 修正正则:仅匹配以.WAV结尾的文件
  recursive = TRUE,
  full.names = TRUE
)
files <- files[grepl("Files for analysis", files)]

# 提取每个文件所属的项目目录(拆分路径后取前6段,对应到Project X层级)
project_paths <- sapply(files, function(x) {
  path_parts <- strsplit(x, "/")[[1]]
  paste(path_parts[1:6], collapse = "/")
})

# 按项目目录分组统计文件数量
project_counts <- table(project_paths)

# 格式化输出为期望样式
for (proj_dir in names(project_counts)) {
  cat(sprintf("%s   %d files\n", proj_dir, project_counts[proj_dir]))
}

关键改进点

  1. 修正正则匹配:将pattern = ".WAV"改为pattern = "\\.WAV$",避免匹配包含WAV字符的非目标文件
  2. 精准过滤文件:通过grepl筛选仅来自Files for analysis目录的文件,彻底排除Backup目录的干扰
  3. 提取项目层级:拆分文件路径并提取到项目根目录,确保分组统计的维度正确
  4. 简化逻辑:去掉冗余的文件夹遍历步骤,直接从文件路径入手统计,提升效率

内容的提问来源于stack exchange,提问作者Lark Davis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 00:50:56