You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言M-K趋势测试循环优化及批量文件命名处理问询

Hey there! Let's tackle your two R scripting problems one by one—they're both perfect candidates for automation, so we'll cut down that repetitive code and streamline your workflow.


1. Optimizing Mann-Kendall Trend Tests with Loops

Your manual code works, but repeating the same lines 60 times is error-prone and tedious. We can use lapply or a for loop to batch-process all your groups, even with varying sample selection rules. Here's a clean, scalable solution:

library(trend)

# Assume your `db` data frame is already loaded
total_groups <- 60  # Adjust to your actual number of groups
group_size <- 10    # Default sample size per group

# --- Step 1: Define group indices (handle custom rules here) ---
# For consecutive 10-sample groups:
group_indices <- lapply(1:total_groups, function(i) {
  start <- (i - 1) * group_size + 1
  end <- i * group_size
  start:end
})

# If you have custom sample rules (e.g., non-consecutive ranges), replace with:
# group_indices <- list(
#   c(1:10), c(15:24), c(30:39),  # Add all 60 custom ranges here
#   ...
# )

# --- Step 2: Run M-K tests in batch ---
mk_results <- lapply(group_indices, function(idx) {
  test_output <- mk.test(db$sum[idx])
  # Extract the exact stats you need, in order
  c(
    test_output$statistic,
    test_output$p.value,
    test_output$estimates[1],
    test_output$estimates[2],
    test_output$estimates[3]
  )
})

# --- Step 3: Convert results to your target data frame format ---
mktests <- as.data.frame(do.call(cbind, mk_results))
colnames(mktests) <- paste0("mk10_", 1:total_groups)
rownames(mktests) <- c("z", "p-value", "S", "varS", "tau")

# Transpose to match your example (rows = test groups, columns = stats)
mktests <- t(mktests)

This code will automatically handle 60 groups (or any number) and adapt to custom sample ranges—just update the group_indices list with your specific selection rules.


2. Batch Processing CSV Files & Extracting IDs

To apply your script to all CSV files in a folder and name outputs with the file's ID, we'll use list.files to grab files, regex to extract IDs, and loop through each file:

library(dplyr)
library(stringr)

# Set your input/output directories (create output folder first!)
input_folder <- "D:/your_csv_directory"
output_folder <- "D:/your_output_directory"

# Get all CSV file paths in the input folder
csv_files <- list.files(input_folder, pattern = "\\.csv$", full.names = TRUE)

# Loop through each file
for (file_path in csv_files) {
  # --- Extract ID from filename ---
  file_name <- basename(file_path)
  # Regex to pull numeric ID (adjust if your IDs have non-numeric characters)
  file_id <- str_extract(file_name, "\\d+")
  
  # --- Load the CSV data ---
  db <- read.csv(file_path)
  
  # --- Run your M-K test code (from problem 1) ---
  # [Insert the M-K loop code here, which generates the `mktests` data frame]
  
  # --- Save results with ID-based filenames ---
  write.csv(mktests, file.path(output_folder, paste0(file_id, "_mk.csv")), row.names = TRUE)
  
  # Example: Save ACF results too (adjust to your actual ACF code)
  acf_output <- acf(db$sum, plot = FALSE)
  acf_df <- as.data.frame(cbind(lag = acf_output$lag, acf = acf_output$acf))
  write.csv(acf_df, file.path(output_folder, paste0(file_id, "_acf.csv")), row.names = FALSE)
}

Notes on ID extraction:

  • If your IDs have a specific pattern (e.g., ID_1234.csv), update the regex to str_extract(file_name, "(?<=ID_)\\d+") to only grab numbers after ID_.
  • Always test the regex on a few filenames first to make sure it captures the correct ID.

内容的提问来源于stack exchange,提问作者Chris

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 18:32:45