变量高效重编码方法、均值按比例计算及批量同规则重编码问询
Hey there! Let's tackle your two data analysis questions using R, since your dataset is structured in R's format. Here's a step-by-step breakdown:
高效变量重编码方法
Below are several practical, efficient methods for recoding variables in R, sorted by flexibility and readability:
dplyr::case_when():The most versatile option for multi-condition recoding, with clean, human-readable syntax that performs well even on large datasets.
Example (convert 1→0, 2→1, keep NA values):library(dplyr) df <- df %>% mutate(var_recoded = case_when( var == 1 ~ 0, var == 2 ~ 1, TRUE ~ NA_real_ ))- Base R
ifelse():Perfect for simple binary recoding, with concise code:df$var_recoded <- ifelse(df$var == 1, 0, ifelse(df$var == 2, 1, NA)) car::recode():A compact tool for quick categorical recoding:library(car) df$var_recoded <- recode(df$var, "1=0; 2=1; else=NA")
按比例计算均值
This typically falls into two common scenarios:
Scenario 1: Weighted Mean (calculate mean using proportional weights)
If you have a weight column (e.g., weight), compute the weighted mean of your target variable:
# Base R weighted_mean <- weighted.mean(df$var, df$weight, na.rm = TRUE) # dplyr approach df %>% summarize(weighted_mean = weighted.mean(var, weight, na.rm = TRUE))
Scenario 2: Grouped Mean with Proportion Distribution
Calculate the mean per group, along with each group's proportion of the total sample:
df %>% group_by(group) %>% summarize( var_mean = mean(var, na.rm = TRUE), group_proportion = n() / nrow(df) )
df_bhs1 Dataset) Your dataset has variables named bhs1_1 to bhs1_10 (all with values 1, 2, or NA). Using dplyr::across() is the most efficient way to recode all these variables at once, without repeating code for each column.
Example: Recode 1→0, 2→1 (with new columns for recoded values)
library(dplyr) df_bhs1 <- df_bhs1 %>% mutate(across(starts_with("bhs1_"), ~ case_when( .x == 1 ~ 0, .x == 2 ~ 1, TRUE ~ NA_real_ ), .names = "{.col}_recoded")) # Adds "_recoded" suffix to new columns
If you want to overwrite the original variables directly:
df_bhs1 <- df_bhs1 %>% mutate(across(starts_with("bhs1_"), ~ case_when( .x == 1 ~ 0, .x == 2 ~ 1, TRUE ~ NA_real_ )))
Alternative: Recode to categorical labels (e.g., "Agree"/"Disagree")
df_bhs1 <- df_bhs1 %>% mutate(across(starts_with("bhs1_"), ~ case_when( .x == 1 ~ "Agree", .x == 2 ~ "Disagree", TRUE ~ NA_character_ ), .names = "{.col}_label"))
Base R alternative (no tidyverse required):
# Select all columns starting with "bhs1_" bhs_cols <- grep("^bhs1_", names(df_bhs1), value = TRUE) # Batch recode and add new columns df_bhs1[paste0(bhs_cols, "_recoded")] <- lapply(df_bhs1[bhs_cols], function(x) { ifelse(x == 1, 0, ifelse(x == 2, 1, NA)) })
内容的提问来源于stack exchange,提问作者Atanas Janackovski

