R语言多变量判断生成新列函数异常求助
Fixing Your R Function for Generating the
treated_psych Column Let's start by breaking down the issues with your current code and then rewrite it to be cleaner, more efficient, and aligned with your actual rules.
What's Wrong With the Original Function?
- Unnecessary (and risky)
assignusage: That loop where you assign each column to a standalone variable is totally unnecessary, and it can cause unexpected variable name conflicts in your environment. You don't need to pull columns out of the data frame to work with them! - Logical error in NA handling: Your first
ifcondition checks if all variables are non-NA before checking for a 1. But your rule says any variable = 1 should set the new column to 1, even if other variables are NA. Right now, if a row has a 1 and an NA, your function incorrectly returns NA instead of 1. - Inefficient row-wise processing: Using
adplyto loop through every row is slow, especially for large datasets. R is designed for vectorized operations—we can do this without looping at all.
Better Implementation (Base R)
First, let's define the list of columns we care about (this makes the code easier to update later):
# List of treatment columns to check tx_columns <- c( "CurrTx6.1_Group", "CurrTx6.1_Ind", "CurrTx6.1_Fam", "CurrTx6.1_Couples", "CurrTx7a_CBTAnx", "CurrTx7b_CBTDep", "CurrTx7c_CBTInsom" )
Now we can generate the treated_psych column in one go using vectorized operations:
dataset$treated_psych <- with(dataset, { # Check if any column in the row is equal to 1 (ignore NAs for this check) has_yes <- rowSums(dataset[tx_columns] == 1, na.rm = TRUE) > 0 # Check if ALL columns in the row are NA all_na <- rowSums(!is.na(dataset[tx_columns])) == 0 # Apply your rules ifelse(all_na, NA_integer_, ifelse(has_yes, 1L, 0L)) })
Better Implementation (tidyverse/dplyr)
If you prefer using the tidyverse, this version is more readable:
library(dplyr) tx_columns <- c( "CurrTx6.1_Group", "CurrTx6.1_Ind", "CurrTx6.1_Fam", "CurrTx6.1_Couples", "CurrTx7a_CBTAnx", "CurrTx7b_CBTDep", "CurrTx7c_CBTInsom" ) dataset <- dataset %>% mutate( # Check for any 1 in the row has_yes = rowSums(across(all_of(tx_columns), ~ .x == 1), na.rm = TRUE) > 0, # Check if all values are NA all_na = rowSums(across(all_of(tx_columns), ~ !is.na(.x))) == 0, # Apply rules with case_when (super readable!) treated_psych = case_when( all_na ~ NA_integer_, has_yes ~ 1L, TRUE ~ 0L ) ) %>% # Optional: Remove the helper columns if you don't need them select(-has_yes, -all_na)
Key Improvements
- Correct logic: Now rows with a 1 and some NAs will correctly return 1, not NA.
- Vectorized speed: No more slow row-wise loops—this will run in a fraction of the time for large datasets.
- Maintainable: If you need to add/remove columns later, you just update the
tx_columnsvector instead of editing multiple lines of code. - Type safety: Using
NA_integer_,1L, and0Lensures the new column stays as integer type (instead of defaulting to numeric).
内容的提问来源于stack exchange,提问作者Katy Torres
相关产品推荐
相关产品推荐

