如何无需安装CRAN R包,从源码统计带示例函数占比?
统计CRAN R包中带帮助页示例的函数占比(无需安装包)
步骤1:提取包内所有函数
从源码的R/目录下,遍历所有.R文件提取函数定义:
# 替换为你的包源码路径 pkg_dir <- "/path/to/your/package" r_files <- list.files(file.path(pkg_dir, "R"), pattern = "\\.R$", full.names = TRUE) all_functions <- c() for (file in r_files) { lines <- readLines(file) # 过滤注释行 lines <- lines[!grepl("^\\s*#", lines)] # 匹配标准函数定义 fun_matches <- regmatches(lines, regexec("^([a-zA-Z0-9._]+)\\s*<-\\s*function\\(", lines)) funs <- sapply(fun_matches, function(x) if (length(x) >= 2) x[2] else NA) all_functions <- c(all_functions, na.omit(funs)) } all_functions <- unique(all_functions) total_fun_count <- length(all_functions)
步骤2:筛选带示例的关联函数
从man/目录的.Rd帮助文件中,找出包含非空\examples{}块的文件,再提取其关联的函数别名:
rd_files <- list.files(file.path(pkg_dir, "man"), pattern = "\\.Rd$", full.names = TRUE) funs_with_examples <- c() for (file in rd_files) { lines <- readLines(file) # 检查是否存在非空的示例块 example_start <- which(grepl("^\\\\examples\\{", lines)) if (length(example_start) > 0) { example_end <- which(grepl("^\\}", lines))[example_end > example_start][1] has_content <- any(grepl("^\\s+.+", lines[(example_start+1):(example_end-1)])) if (has_content) { # 提取该帮助页关联的所有函数别名 alias_matches <- regmatches(lines, regexec("^\\\\alias\\{([^}]+)\\}", lines)) aliases <- sapply(alias_matches, function(x) if (length(x) >= 2) x[2] else NA) funs_with_examples <- c(funs_with_examples, na.omit(aliases)) } } } funs_with_examples <- unique(funs_with_examples) # 只保留真实存在的函数(排除数据集等非函数别名) funs_with_examples <- intersect(funs_with_examples, all_functions) has_example_count <- length(funs_with_examples)
步骤3:计算占比
直接计算并输出结果:
proportion <- has_example_count / total_fun_count cat(sprintf("带帮助页示例的函数占比:%.2f%%\n", proportion * 100))
注意事项
- 正则匹配仅覆盖标准命名的函数,若包内有特殊命名的函数,需调整正则规则。
- 会自动过滤空示例块(仅含
\examples{}大括号无代码的情况),以及帮助页中关联的数据集别名。
内容的提问来源于stack exchange,提问作者Patrick
相关产品推荐
相关产品推荐

