You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言多变量判断生成新列函数异常求助

Fixing Your R Function for Generating the treated_psych Column

Let's start by breaking down the issues with your current code and then rewrite it to be cleaner, more efficient, and aligned with your actual rules.

What's Wrong With the Original Function?

  1. Unnecessary (and risky) assign usage: That loop where you assign each column to a standalone variable is totally unnecessary, and it can cause unexpected variable name conflicts in your environment. You don't need to pull columns out of the data frame to work with them!
  2. Logical error in NA handling: Your first if condition checks if all variables are non-NA before checking for a 1. But your rule says any variable = 1 should set the new column to 1, even if other variables are NA. Right now, if a row has a 1 and an NA, your function incorrectly returns NA instead of 1.
  3. Inefficient row-wise processing: Using adply to loop through every row is slow, especially for large datasets. R is designed for vectorized operations—we can do this without looping at all.

Better Implementation (Base R)

First, let's define the list of columns we care about (this makes the code easier to update later):

# List of treatment columns to check
tx_columns <- c(
  "CurrTx6.1_Group", "CurrTx6.1_Ind", "CurrTx6.1_Fam",
  "CurrTx6.1_Couples", "CurrTx7a_CBTAnx", "CurrTx7b_CBTDep",
  "CurrTx7c_CBTInsom"
)

Now we can generate the treated_psych column in one go using vectorized operations:

dataset$treated_psych <- with(dataset, {
  # Check if any column in the row is equal to 1 (ignore NAs for this check)
  has_yes <- rowSums(dataset[tx_columns] == 1, na.rm = TRUE) > 0
  # Check if ALL columns in the row are NA
  all_na <- rowSums(!is.na(dataset[tx_columns])) == 0
  
  # Apply your rules
  ifelse(all_na, NA_integer_, ifelse(has_yes, 1L, 0L))
})

Better Implementation (tidyverse/dplyr)

If you prefer using the tidyverse, this version is more readable:

library(dplyr)

tx_columns <- c(
  "CurrTx6.1_Group", "CurrTx6.1_Ind", "CurrTx6.1_Fam",
  "CurrTx6.1_Couples", "CurrTx7a_CBTAnx", "CurrTx7b_CBTDep",
  "CurrTx7c_CBTInsom"
)

dataset <- dataset %>%
  mutate(
    # Check for any 1 in the row
    has_yes = rowSums(across(all_of(tx_columns), ~ .x == 1), na.rm = TRUE) > 0,
    # Check if all values are NA
    all_na = rowSums(across(all_of(tx_columns), ~ !is.na(.x))) == 0,
    # Apply rules with case_when (super readable!)
    treated_psych = case_when(
      all_na ~ NA_integer_,
      has_yes ~ 1L,
      TRUE ~ 0L
    )
  ) %>%
  # Optional: Remove the helper columns if you don't need them
  select(-has_yes, -all_na)

Key Improvements

  • Correct logic: Now rows with a 1 and some NAs will correctly return 1, not NA.
  • Vectorized speed: No more slow row-wise loops—this will run in a fraction of the time for large datasets.
  • Maintainable: If you need to add/remove columns later, you just update the tx_columns vector instead of editing multiple lines of code.
  • Type safety: Using NA_integer_, 1L, and 0L ensures the new column stays as integer type (instead of defaulting to numeric).

内容的提问来源于stack exchange,提问作者Katy Torres

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 03:47:48