You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何统计R包的代码行数?支持本地源码与tar.gz包两种输入

统计R包有效代码行数(含注释行统计)

下面是一个可直接在R环境运行的函数,支持处理本地R源码文件夹或tar.gz格式的R包文件,能准确识别R注释(排除字符串中的#),同时统计有效代码行、注释行和空行数量,且效率较高:

count_r_lines <- function(path) {
  # 判断输入类型:文件夹或tar.gz包
  if (file.exists(path)) {
    if (dir.exists(path)) {
      # 本地R文件夹,获取所有.R文件
      r_files <- list.files(path, pattern = "\\.R$", full.names = TRUE, recursive = TRUE)
    } else if (grepl("\\.tar\\.gz$", path)) {
      # 临时目录解压tar.gz包
      temp_dir <- tempfile()
      dir.create(temp_dir)
      untar(path, exdir = temp_dir)
      # 定位包内的R文件夹
      pkg_root <- list.dirs(temp_dir, recursive = FALSE)[1]
      r_files <- list.files(file.path(pkg_root, "R"), pattern = "\\.R$", full.names = TRUE)
      # 自动清理临时目录
      on.exit(unlink(temp_dir, recursive = TRUE))
    } else {
      stop("输入路径不是R文件夹或tar.gz格式的包文件")
    }
  } else {
    stop("输入路径不存在")
  }

  if (length(r_files) == 0) {
    stop("未找到任何.R文件")
  }

  # 辅助函数:判断单行类型(空行/注释行/代码行)
  classify_line <- function(line) {
    line_trim <- trimws(line)
    if (line_trim == "") {
      return("empty")
    }
    # 移除配对引号内的内容,避免误判字符串中的#
    line_no_str <- gsub('"([^"\\\\]|\\\\.)*"', '', line_trim)
    line_no_str <- gsub("'([^'\\\\]|\\\\.)*'", '', line_no_str)
    line_no_str_trim <- trimws(line_no_str)
    if (startsWith(line_no_str_trim, "#")) {
      return("comment")
    } else {
      return("code")
    }
  }

  # 统计所有文件的行数
  total_counts <- list(code = 0, comment = 0, empty = 0)
  for (file in r_files) {
    lines <- readLines(file, warn = FALSE)
    line_types <- sapply(lines, classify_line, USE.NAMES = FALSE)
    counts <- table(line_types)
    # 合并统计结果
    total_counts$code <- total_counts$code + ifelse("code" %in% names(counts), counts["code"], 0)
    total_counts$comment <- total_counts$comment + ifelse("comment" %in% names(counts), counts["comment"], 0)
    total_counts$empty <- total_counts$empty + ifelse("empty" %in% names(counts), counts["empty"], 0)
  }

  # 打印结果并返回
  cat(sprintf("有效代码行数: %d\n", total_counts$code))
  cat(sprintf("注释行数: %d\n", total_counts$comment))
  cat(sprintf("空行数: %d\n", total_counts$empty))
  return(invisible(total_counts))
}

使用方法

  • 处理本地自研R包的R文件夹:
    count_r_lines("path/to/your/package/R")
    
  • 处理tar.gz格式的包文件:
    count_r_lines("path/to/your/package.tar.gz")
    

说明

  1. 准确性:通过正则移除字符串内的内容,避免将字符串中的#误判为注释,同时准确识别行首带空格的整行注释。
  2. 效率:基于base R原生函数实现,无额外包依赖,采用线性遍历逻辑,耗时可控。
  3. 输出:控制台直接打印三类行数统计,同时返回包含统计结果的列表(可通过变量接收后续处理)。

内容的提问来源于stack exchange,提问作者Tripartio

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 05:18:27