You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

编写R脚本处理多CSV文件:分类变量计数与连续变量均值计算

Hey there! Let's tackle your R scripting and batch processing needs step by step. I'll break this into two focused R scripts plus a shell wrapper to handle all 60 CSV files seamlessly. Here's what you need:

1. Reusable Mean Calculation Script (mean_calculator.R)

This script defines a flexible function to calculate grouped means for continuous variables—you can use it standalone or call it from other scripts. It includes input validation to avoid common errors:

# mean_calculator.R: 计算数据框中连续变量的分组均值
library(dplyr)

calculate_grouped_means <- function(data_df, group_col, continuous_cols) {
  # 检查指定列是否存在于数据框中
  if (!all(c(group_col, continuous_cols) %in% colnames(data_df))) {
    stop("Error: One or more specified columns don't exist in the data frame!")
  }
  
  # 按分类变量分组,计算连续变量的均值(忽略NA值)
  grouped_means <- data_df %>%
    group_by({{ group_col }}) %>%
    summarise(across({{ continuous_cols }}, mean, na.rm = TRUE), .groups = "drop")
  
  return(grouped_means)
}

# 交互式运行时的测试代码(方便调试)
if (interactive()) {
  # 生成测试数据
  test_data <- data.frame(
    category = rep(c("X", "Y", "Z"), each = 15),
    temp = rnorm(45, 20, 3),
    pressure = rnorm(45, 1013, 5),
    humidity = rnorm(45, 60, 10)
  )
  
  # 调用函数计算均值
  result <- calculate_grouped_means(test_data, "category", c("temp", "pressure", "humidity"))
  print("测试结果:")
  print(result)
}

Quick Notes:

  • Make sure you have the dplyr package installed (install.packages("dplyr")).
  • The {{ }} syntax lets you pass column names as plain text without extra quoting.
2. Single CSV Processing Script (process_single_csv.R)

This script is built to handle one CSV file at a time—perfect for being called in a shell loop. It will:

  • Count value frequencies for your 100 categorical variables
  • Calculate grouped means for the 3 continuous variables (grouped by each categorical variable's values)
  • Save results to a specified output directory
# process_single_csv.R: 处理单个CSV文件,输出分类变量计数和连续变量分组均值
library(dplyr)
library(tools)

# 从命令行获取输入参数:第一个是CSV文件路径,第二个是输出目录
args <- commandArgs(trailingOnly = TRUE)
if (length(args) != 2) {
  stop("Usage: Rscript process_single_csv.R <input_csv_path> <output_directory>")
}

input_file <- args[1]
output_dir <- args[2]

# 创建输出目录(如果不存在)
if (!dir.exists(output_dir)) {
  dir.create(output_dir, recursive = TRUE)
}

# 读取CSV数据(根据你的数据格式调整read.csv参数,比如sep, header等)
data <- read.csv(input_file, stringsAsFactors = FALSE)

# --------------------------
# 配置:根据你的实际列名调整这里!
# 假设前100列是数值型分类变量,最后3列是连续变量
category_cols <- colnames(data)[1:100]
continuous_cols <- colnames(data)[101:103]
# 如果列名有特定前缀(比如cat_1到cat_100),可以用:
# category_cols <- grep("^cat_", colnames(data), value = TRUE)
# continuous_cols <- grep("^cont_", colnames(data), value = TRUE)
# --------------------------

# 1. 统计每个分类变量的取值计数
category_counts <- lapply(category_cols, function(col) {
  data %>%
    count({{ col }}, name = "count") %>%
    rename(category_value = {{ col }}) %>%
    mutate(variable_name = col)
})
category_counts <- bind_rows(category_counts)

# 2. 按分类变量取值计算连续变量的均值
grouped_means <- lapply(category_cols, function(col) {
  data %>%
    group_by({{ col }}) %>%
    summarise(across(continuous_cols, mean, na.rm = TRUE), .groups = "drop") %>%
    rename(category_value = {{ col }}) %>%
    mutate(variable_name = col)
})
grouped_means <- bind_rows(grouped_means)

# 保存结果到输出目录(用原文件名作为前缀)
file_prefix <- file_path_sans_ext(basename(input_file))
write.csv(category_counts, file.path(output_dir, paste0(file_prefix, "_category_counts.csv")), row.names = FALSE)
write.csv(grouped_means, file.path(output_dir, paste0(file_prefix, "_grouped_means.csv")), row.names = FALSE)

# 打印处理完成提示
cat(sprintf("Successfully processed: %s\nResults saved to: %s\n", input_file, output_dir))
3. Shell Script for Batch Processing (batch_process.sh)

This shell script will loop through all CSV files in your input directory and call the R script for each one. Save it as batch_process.sh:

#!/bin/bash

# 配置你的输入和输出目录
INPUT_DIR="./test_subjects"  # 存放60个CSV文件的文件夹
OUTPUT_DIR="./batch_results" # 结果输出文件夹

# 检查输入目录是否存在
if [ ! -d "$INPUT_DIR" ]; then
    echo "Error: Input directory $INPUT_DIR does not exist!"
    exit 1
fi

# 创建输出目录
mkdir -p "$OUTPUT_DIR"

# 循环处理每个CSV文件
for csv_file in "$INPUT_DIR"/*.csv; do
    # 跳过非文件(比如目录下没有CSV文件时的通配符本身)
    [ -f "$csv_file" ] || continue
    
    echo "Processing file: $csv_file"
    # 调用R脚本处理单个文件
    Rscript process_single_csv.R "$csv_file" "$OUTPUT_DIR"
done

echo "All files processed! Results are in $OUTPUT_DIR"

How to Run the Batch Process:

  1. Make the shell script executable:
    chmod +x batch_process.sh
    
  2. Run it:
    ./batch_process.sh
    

Final Tips:

  • Test with one CSV file first using Rscript process_single_csv.R path/to/test.csv path/to/test_output to make sure everything works.
  • Adjust the column selection logic in process_single_csv.R to match your actual CSV structure.
  • If your categorical variables have NA values, add na.rm = TRUE to the count() function if you want to exclude them.

内容的提问来源于stack exchange,提问作者Sam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:58:03