You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

变量高效重编码方法、均值按比例计算及批量同规则重编码问询

Hey there! Let's tackle your two data analysis questions using R, since your dataset is structured in R's format. Here's a step-by-step breakdown:

1. 高效的变量重编码方法 & 按比例计算均值

高效变量重编码方法

Below are several practical, efficient methods for recoding variables in R, sorted by flexibility and readability:

  • dplyr::case_when():The most versatile option for multi-condition recoding, with clean, human-readable syntax that performs well even on large datasets.
    Example (convert 1→0, 2→1, keep NA values):
    library(dplyr)
    df <- df %>% mutate(var_recoded = case_when(
      var == 1 ~ 0,
      var == 2 ~ 1,
      TRUE ~ NA_real_
    ))
    
  • Base R ifelse():Perfect for simple binary recoding, with concise code:
    df$var_recoded <- ifelse(df$var == 1, 0, ifelse(df$var == 2, 1, NA))
    
  • car::recode():A compact tool for quick categorical recoding:
    library(car)
    df$var_recoded <- recode(df$var, "1=0; 2=1; else=NA")
    

按比例计算均值

This typically falls into two common scenarios:

Scenario 1: Weighted Mean (calculate mean using proportional weights)

If you have a weight column (e.g., weight), compute the weighted mean of your target variable:

# Base R
weighted_mean <- weighted.mean(df$var, df$weight, na.rm = TRUE)

# dplyr approach
df %>% summarize(weighted_mean = weighted.mean(var, weight, na.rm = TRUE))

Scenario 2: Grouped Mean with Proportion Distribution

Calculate the mean per group, along with each group's proportion of the total sample:

df %>%
  group_by(group) %>%
  summarize(
    var_mean = mean(var, na.rm = TRUE),
    group_proportion = n() / nrow(df)
  )

2. Batch Recode Multiple Variables (For Your df_bhs1 Dataset)

Your dataset has variables named bhs1_1 to bhs1_10 (all with values 1, 2, or NA). Using dplyr::across() is the most efficient way to recode all these variables at once, without repeating code for each column.

Example: Recode 1→0, 2→1 (with new columns for recoded values)

library(dplyr)

df_bhs1 <- df_bhs1 %>%
  mutate(across(starts_with("bhs1_"), ~ case_when(
    .x == 1 ~ 0,
    .x == 2 ~ 1,
    TRUE ~ NA_real_
  ), .names = "{.col}_recoded")) # Adds "_recoded" suffix to new columns

If you want to overwrite the original variables directly:

df_bhs1 <- df_bhs1 %>%
  mutate(across(starts_with("bhs1_"), ~ case_when(
    .x == 1 ~ 0,
    .x == 2 ~ 1,
    TRUE ~ NA_real_
  )))

Alternative: Recode to categorical labels (e.g., "Agree"/"Disagree")

df_bhs1 <- df_bhs1 %>%
  mutate(across(starts_with("bhs1_"), ~ case_when(
    .x == 1 ~ "Agree",
    .x == 2 ~ "Disagree",
    TRUE ~ NA_character_
  ), .names = "{.col}_label"))

Base R alternative (no tidyverse required):

# Select all columns starting with "bhs1_"
bhs_cols <- grep("^bhs1_", names(df_bhs1), value = TRUE)
# Batch recode and add new columns
df_bhs1[paste0(bhs_cols, "_recoded")] <- lapply(df_bhs1[bhs_cols], function(x) {
  ifelse(x == 1, 0, ifelse(x == 2, 1, NA))
})

内容的提问来源于stack exchange,提问作者Atanas Janackovski

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:57:37