编写R脚本处理多CSV文件:分类变量计数与连续变量均值计算
Hey there! Let's tackle your R scripting and batch processing needs step by step. I'll break this into two focused R scripts plus a shell wrapper to handle all 60 CSV files seamlessly. Here's what you need:
1. Reusable Mean Calculation Script (
mean_calculator.R) This script defines a flexible function to calculate grouped means for continuous variables—you can use it standalone or call it from other scripts. It includes input validation to avoid common errors:
# mean_calculator.R: 计算数据框中连续变量的分组均值 library(dplyr) calculate_grouped_means <- function(data_df, group_col, continuous_cols) { # 检查指定列是否存在于数据框中 if (!all(c(group_col, continuous_cols) %in% colnames(data_df))) { stop("Error: One or more specified columns don't exist in the data frame!") } # 按分类变量分组,计算连续变量的均值(忽略NA值) grouped_means <- data_df %>% group_by({{ group_col }}) %>% summarise(across({{ continuous_cols }}, mean, na.rm = TRUE), .groups = "drop") return(grouped_means) } # 交互式运行时的测试代码(方便调试) if (interactive()) { # 生成测试数据 test_data <- data.frame( category = rep(c("X", "Y", "Z"), each = 15), temp = rnorm(45, 20, 3), pressure = rnorm(45, 1013, 5), humidity = rnorm(45, 60, 10) ) # 调用函数计算均值 result <- calculate_grouped_means(test_data, "category", c("temp", "pressure", "humidity")) print("测试结果:") print(result) }
Quick Notes:
- Make sure you have the
dplyrpackage installed (install.packages("dplyr")). - The
{{ }}syntax lets you pass column names as plain text without extra quoting.
2. Single CSV Processing Script (
process_single_csv.R) This script is built to handle one CSV file at a time—perfect for being called in a shell loop. It will:
- Count value frequencies for your 100 categorical variables
- Calculate grouped means for the 3 continuous variables (grouped by each categorical variable's values)
- Save results to a specified output directory
# process_single_csv.R: 处理单个CSV文件,输出分类变量计数和连续变量分组均值 library(dplyr) library(tools) # 从命令行获取输入参数:第一个是CSV文件路径,第二个是输出目录 args <- commandArgs(trailingOnly = TRUE) if (length(args) != 2) { stop("Usage: Rscript process_single_csv.R <input_csv_path> <output_directory>") } input_file <- args[1] output_dir <- args[2] # 创建输出目录(如果不存在) if (!dir.exists(output_dir)) { dir.create(output_dir, recursive = TRUE) } # 读取CSV数据(根据你的数据格式调整read.csv参数,比如sep, header等) data <- read.csv(input_file, stringsAsFactors = FALSE) # -------------------------- # 配置:根据你的实际列名调整这里! # 假设前100列是数值型分类变量,最后3列是连续变量 category_cols <- colnames(data)[1:100] continuous_cols <- colnames(data)[101:103] # 如果列名有特定前缀(比如cat_1到cat_100),可以用: # category_cols <- grep("^cat_", colnames(data), value = TRUE) # continuous_cols <- grep("^cont_", colnames(data), value = TRUE) # -------------------------- # 1. 统计每个分类变量的取值计数 category_counts <- lapply(category_cols, function(col) { data %>% count({{ col }}, name = "count") %>% rename(category_value = {{ col }}) %>% mutate(variable_name = col) }) category_counts <- bind_rows(category_counts) # 2. 按分类变量取值计算连续变量的均值 grouped_means <- lapply(category_cols, function(col) { data %>% group_by({{ col }}) %>% summarise(across(continuous_cols, mean, na.rm = TRUE), .groups = "drop") %>% rename(category_value = {{ col }}) %>% mutate(variable_name = col) }) grouped_means <- bind_rows(grouped_means) # 保存结果到输出目录(用原文件名作为前缀) file_prefix <- file_path_sans_ext(basename(input_file)) write.csv(category_counts, file.path(output_dir, paste0(file_prefix, "_category_counts.csv")), row.names = FALSE) write.csv(grouped_means, file.path(output_dir, paste0(file_prefix, "_grouped_means.csv")), row.names = FALSE) # 打印处理完成提示 cat(sprintf("Successfully processed: %s\nResults saved to: %s\n", input_file, output_dir))
3. Shell Script for Batch Processing (
batch_process.sh) This shell script will loop through all CSV files in your input directory and call the R script for each one. Save it as batch_process.sh:
#!/bin/bash # 配置你的输入和输出目录 INPUT_DIR="./test_subjects" # 存放60个CSV文件的文件夹 OUTPUT_DIR="./batch_results" # 结果输出文件夹 # 检查输入目录是否存在 if [ ! -d "$INPUT_DIR" ]; then echo "Error: Input directory $INPUT_DIR does not exist!" exit 1 fi # 创建输出目录 mkdir -p "$OUTPUT_DIR" # 循环处理每个CSV文件 for csv_file in "$INPUT_DIR"/*.csv; do # 跳过非文件(比如目录下没有CSV文件时的通配符本身) [ -f "$csv_file" ] || continue echo "Processing file: $csv_file" # 调用R脚本处理单个文件 Rscript process_single_csv.R "$csv_file" "$OUTPUT_DIR" done echo "All files processed! Results are in $OUTPUT_DIR"
How to Run the Batch Process:
- Make the shell script executable:
chmod +x batch_process.sh - Run it:
./batch_process.sh
Final Tips:
- Test with one CSV file first using
Rscript process_single_csv.R path/to/test.csv path/to/test_outputto make sure everything works. - Adjust the column selection logic in
process_single_csv.Rto match your actual CSV structure. - If your categorical variables have NA values, add
na.rm = TRUEto thecount()function if you want to exclude them.
内容的提问来源于stack exchange,提问作者Sam
相关产品推荐
相关产品推荐

