字符串替换解析及CSV计算规则转R语言并验证的技术咨询
Got it, let's walk through both of your tasks with practical, actionable R code—this should cover everything you need for string replacement and CSV data validation.
If you need to swap placeholders in a text template with dynamic variables, two straightforward approaches work great in R: using the glue package (for readable natural language templates) or base R's sprintf (for strict formatting needs).
Example with glue (most intuitive for text)
# Install and load the glue package if you haven't already if (!require(glue)) install.packages("glue") library(glue) # Define your text template with {variable_name} placeholders text_template <- "Task update: {task_name} assigned to {assignee} is {status} (due {due_date})." # Define your replacement variables task_name <- "CSV Data Validation" assignee <- "Sven" status <- "in progress" due_date <- "2024-06-30" # Run replacement and parsing parsed_text <- glue(text_template) print(parsed_text)
Example with base R sprintf (great for numeric formatting)
# Template uses %s for strings, %d for integers, %f for decimals numeric_template <- "Calculation result: The average of %d values is %.2f." # Variables value_count <- 150 average_value <- 78.456 # Parse and replace parsed_numeric <- sprintf(numeric_template, value_count, average_value) print(parsed_numeric)
Let's break this into three clear steps, as you outlined: translating rules to R, importing your files, and applying validation.
2.1 Translate Your Calculation Rules to R Logic
First, convert your specific rules into vectorized R code (R works best with vector operations instead of loops). For example:
- If your rule is "Mark rows where Column X > 100 AND Column Y = 'Approved'", translate it to:
# For a single data frame df$condition_met <- ifelse(df$X > 100 & df$Y == "Approved", TRUE, FALSE) - Adjust this to match your actual rule (e.g., row sums, column ratios, etc.)—vectorized logic will scale to all your worksheets.
2.2 Import Multiple CSV Files/Worksheets into R
Assuming your files are named like C 01.00.csv, F 08.01.b.csv and stored in a single folder, use the tidyverse to read them into a named list (with clean names like C0100, F0801b as you requested):
# Install and load tidyverse if needed if (!require(tidyverse)) install.packages("tidyverse") library(tidyverse) # Set the path to your CSV folder (replace with your actual path) csv_folder <- "./your_csv_directory/" # Get all matching CSV files and read them into a named list csv_worksheets <- list.files( path = csv_folder, pattern = "^(C|F) .*\\.csv$", # Match files starting with C/F followed by space full.names = TRUE ) %>% # Clean up filenames to create list names (remove spaces, dots, .csv) set_names(str_remove_all(basename(.), "\\.csv| |\\.")) %>% # Read each CSV into a data frame map(read_csv) # Access a specific worksheet matrix (convert data frame to matrix if needed) F0801b_matrix <- as.matrix(csv_worksheets$F0801b)
Note: If you're working with Excel worksheets (not separate CSVs), use readxl::read_excel with the sheet parameter instead of read_csv.
2.3 Apply Validation Rules to All Worksheets
Create a reusable function for your rule, then apply it to every worksheet in your list:
# Define your validation rule function (adjust this to match your actual rule) validate_worksheet <- function(data_matrix) { # Example rule: Return TRUE if row sum is greater than 500 (ignore missing values) row_totals <- rowSums(data_matrix, na.rm = TRUE) return(row_totals > 500) } # Apply the rule to all worksheets and store results validation_outcomes <- csv_worksheets %>% map(~ validate_worksheet(as.matrix(.x))) # View results for a specific worksheet print(validation_outcomes$F0801b) # Optional: Add validation results back to the original data frames processed_worksheets <- csv_worksheets %>% imap(~ mutate(.x, condition_met = validate_worksheet(as.matrix(.x)))) # Save processed data to new CSV files processed_worksheets %>% iwalk(~ write_csv(.x, file.path(csv_folder, paste0(.y, "_processed.csv"))))
If your rule is more complex (e.g., cross-column calculations, group-wise checks), just tweak the validate_worksheet function to match—R's vectorized operations make this easy to scale.
内容的提问来源于stack exchange,提问作者Sven

