如何在R语言中基于另外两列的特定行值创建新变量?
Got it, let's work through this problem step by step. First, I'll assume your dataset has a structure like this (since each patient has 4 rows, it's almost certainly split by ear and time point):
| patient_id | ear | time_point | chemo_dose |
|---|---|---|---|
| 1 | L | chemo dose 1 | 100 |
| 1 | R | chemo dose 1 | 120 |
| 1 | L | chemo dose 2 | 150 |
| 1 | R | chemo dose 2 | 180 |
| 2 | L | chemo dose 1 | 90 |
| 2 | R | chemo dose 1 | 110 |
| 2 | L | chemo dose 2 | 140 |
| 2 | R | chemo dose 2 | 170 |
Your goal is to create a new variable that pulls chemo_dose from:
- The first row of "chemo dose 1" per patient (likely the left ear row)
- The second row of "chemo dose 2" per patient (likely the right ear row)
Here are two reliable ways to implement this in R:
Method 1: Using dplyr (Tidyverse)
This is the most readable approach, especially if you're already working with tidyverse tools. We'll group by patient and target the exact rows we need.
Option A: Target rows by ear (most reliable)
If you have an ear column marking left/right, use this instead of relying on row order (which can shift unexpectedly):
library(dplyr) df_new <- df %>% group_by(patient_id) %>% mutate( new_chemo_var = case_when( # Pull value from left ear during chemo dose 1 time_point == "chemo dose 1" & ear == "L" ~ chemo_dose, # Pull value from right ear during chemo dose 2 time_point == "chemo dose 2" & ear == "R" ~ chemo_dose, # Set all other rows to NA (adjust this if you need a different default) TRUE ~ NA_real_ ) ) %>% ungroup()
Option B: Target rows by position (if row order is guaranteed)
If you're 100% sure each patient's rows follow the order: left dose1 → right dose1 → left dose2 → right dose2, you can use row numbers within each patient group:
df_new <- df %>% group_by(patient_id) %>% mutate( new_chemo_var = case_when( # 1st row in group = left ear dose1 row_number() == 1 ~ chemo_dose, # 4th row in group = right ear dose2 row_number() == 4 ~ chemo_dose, TRUE ~ NA_real_ ) ) %>% ungroup()
Method 2: Using Base R
If you prefer not to load additional libraries, here's a base R approach using splitting and looping:
# Initialize the new variable with NA values df$new_chemo_var <- NA_real_ # Split the dataset into patient-specific groups patient_groups <- split(df, df$patient_id) # Iterate over each patient group to assign values for (group_name in names(patient_groups)) { group <- patient_groups[[group_name]] # Find indices for the rows we need dose1_left_idx <- which(group$time_point == "chemo dose 1" & group$ear == "L") dose2_right_idx <- which(group$time_point == "chemo dose 2" & group$ear == "R") # Assign the chemo dose values to the new variable if (length(dose1_left_idx) > 0) { group$new_chemo_var[dose1_left_idx] <- group$chemo_dose[dose1_left_idx] } if (length(dose2_right_idx) > 0) { group$new_chemo_var[dose2_right_idx] <- group$chemo_dose[dose2_right_idx] } # Update the group in the list patient_groups[[group_name]] <- group } # Combine all groups back into a single data frame df_new <- do.call(rbind, patient_groups)
Quick Tips
- Always prioritize using explicit identifiers like
earover row order—row order can change during data cleaning or importing, leading to errors. - If you want to propagate these values to all rows for each patient (instead of just the target rows), you can use
first()/last()within the group to pull the value and assign it to every row.
内容的提问来源于stack exchange,提问作者sabc04

