基于用户输入动态处理R Data Frame空值并新增标记列的需求
Got it, let's tackle this problem step by step. Here's a robust solution that meets both of your requirements, with explanations for each part:
Solution Code (using dplyr for readability)
First, we'll use the dplyr package since it makes data manipulation more intuitive, but I'll also include a base R alternative below.
# Load the dplyr library (install first if needed: install.packages("dplyr")) library(dplyr) # Define your original data frame df <- data.frame( HOD_ID = c("","KTN2252","ANA2719","ITI2624","DEV2698","HRT2921","","KTN2624","ANA2548","ITI2535","DEV2732","HRT2837","ERV2951","KTN2542","ANA2813","ITI2210"), DEPT_Head = c("DEV0001","KTN2252","ANA2719","ITI2624","DEV2698","HRT2921","ERV0000","KTN2624","ANA2548","ITI2535","DEV2732","HRT2837","ERV2951","KTN2542","ANA2813","ITI2210"), stringsAsFactors = FALSE # Critical to handle empty strings correctly ) # Function to process the data with dynamic DEPT_Head input process_dept_data <- function(input_df, target_dept_head) { # Step 1: Validate the target DEPT_Head has an empty HOD_ID target_row <- input_df %>% filter(DEPT_Head == target_dept_head) # Throw an error if validation fails stopifnot( nrow(target_row) == 1, # Ensure the target exists exactly once target_row$HOD_ID == "" # Ensure its HOD_ID is empty ) # Step 2: Add the if_blank column processed_df <- input_df %>% mutate( if_blank = case_when( # Assign 1 only to non-target rows with empty HOD_ID DEPT_Head != target_dept_head & HOD_ID == "" ~ 1, # Assign 0 to all other rows (adjust to NA if preferred) TRUE ~ 0 ) ) return(processed_df) } # Example usage with your sample input: DEPT_Head = "DEV0001" result <- process_dept_data(df, "DEV0001") print(result)
Step-by-Step Explanation
- Data Frame Setup: We set
stringsAsFactors = FALSEto avoid unexpected behavior with empty strings (factors can treat blanks as a separate level, which we don't want here). - Validation Check: The function first finds the row matching your input
target_dept_head. We usestopifnotto enforce two rules:- The target DEPT_Head exists exactly once in the data.
- The corresponding
HOD_IDis empty. If either rule is broken, the function throws a clear error to help you catch issues early.
- Adding
if_blank: Usingmutateandcase_when, we precisely set the value ofif_blank:- 1 for rows that aren't the target DEPT_Head and have an empty HOD_ID.
- 0 for all other rows (you can change this to
NA_integer_if you prefer missing values instead of 0).
Base R Alternative (no external packages)
If you don't want to use dplyr, here's an equivalent solution using only base R functions:
process_dept_data_base <- function(input_df, target_dept_head) { # Validation step target_idx <- which(input_df$DEPT_Head == target_dept_head) stopifnot( length(target_idx) == 1, input_df$HOD_ID[target_idx] == "" ) # Initialize if_blank column to 0 input_df$if_blank <- 0 # Set if_blank to 1 for qualifying rows input_df$if_blank[input_df$DEPT_Head != target_dept_head & input_df$HOD_ID == ""] <- 1 return(input_df) } # Example usage result_base <- process_dept_data_base(df, "DEV0001") print(result_base)
内容的提问来源于stack exchange,提问作者newcomer
相关产品推荐
相关产品推荐

