R语言M-K趋势测试循环优化及批量文件命名处理问询
Hey there! Let's tackle your two R scripting problems one by one—they're both perfect candidates for automation, so we'll cut down that repetitive code and streamline your workflow.
Your manual code works, but repeating the same lines 60 times is error-prone and tedious. We can use lapply or a for loop to batch-process all your groups, even with varying sample selection rules. Here's a clean, scalable solution:
library(trend) # Assume your `db` data frame is already loaded total_groups <- 60 # Adjust to your actual number of groups group_size <- 10 # Default sample size per group # --- Step 1: Define group indices (handle custom rules here) --- # For consecutive 10-sample groups: group_indices <- lapply(1:total_groups, function(i) { start <- (i - 1) * group_size + 1 end <- i * group_size start:end }) # If you have custom sample rules (e.g., non-consecutive ranges), replace with: # group_indices <- list( # c(1:10), c(15:24), c(30:39), # Add all 60 custom ranges here # ... # ) # --- Step 2: Run M-K tests in batch --- mk_results <- lapply(group_indices, function(idx) { test_output <- mk.test(db$sum[idx]) # Extract the exact stats you need, in order c( test_output$statistic, test_output$p.value, test_output$estimates[1], test_output$estimates[2], test_output$estimates[3] ) }) # --- Step 3: Convert results to your target data frame format --- mktests <- as.data.frame(do.call(cbind, mk_results)) colnames(mktests) <- paste0("mk10_", 1:total_groups) rownames(mktests) <- c("z", "p-value", "S", "varS", "tau") # Transpose to match your example (rows = test groups, columns = stats) mktests <- t(mktests)
This code will automatically handle 60 groups (or any number) and adapt to custom sample ranges—just update the group_indices list with your specific selection rules.
To apply your script to all CSV files in a folder and name outputs with the file's ID, we'll use list.files to grab files, regex to extract IDs, and loop through each file:
library(dplyr) library(stringr) # Set your input/output directories (create output folder first!) input_folder <- "D:/your_csv_directory" output_folder <- "D:/your_output_directory" # Get all CSV file paths in the input folder csv_files <- list.files(input_folder, pattern = "\\.csv$", full.names = TRUE) # Loop through each file for (file_path in csv_files) { # --- Extract ID from filename --- file_name <- basename(file_path) # Regex to pull numeric ID (adjust if your IDs have non-numeric characters) file_id <- str_extract(file_name, "\\d+") # --- Load the CSV data --- db <- read.csv(file_path) # --- Run your M-K test code (from problem 1) --- # [Insert the M-K loop code here, which generates the `mktests` data frame] # --- Save results with ID-based filenames --- write.csv(mktests, file.path(output_folder, paste0(file_id, "_mk.csv")), row.names = TRUE) # Example: Save ACF results too (adjust to your actual ACF code) acf_output <- acf(db$sum, plot = FALSE) acf_df <- as.data.frame(cbind(lag = acf_output$lag, acf = acf_output$acf)) write.csv(acf_df, file.path(output_folder, paste0(file_id, "_acf.csv")), row.names = FALSE) }
Notes on ID extraction:
- If your IDs have a specific pattern (e.g.,
ID_1234.csv), update the regex tostr_extract(file_name, "(?<=ID_)\\d+")to only grab numbers afterID_. - Always test the regex on a few filenames first to make sure it captures the correct ID.
内容的提问来源于stack exchange,提问作者Chris

