基于CSES数据集批量构建ideology_voted变量的技术问询
ideology_voted变量构建问题 Hey there! Let's work through this problem together. You need to create a variable ideology_voted that maps each respondent's voted party to its corresponding self-perceived ideology—and you want to avoid writing repetitive conditional code for all 9 parties (A-I). Let's break this down into two parts: first fixing the repeated conditional logic you're struggling with, then scaling it up for all parties.
1. 先解决重复条件的基础写法
If you're just starting out, using dplyr::case_when() makes the conditional logic clear and readable, even for a few parties. Here's how to apply it to your sample data:
library(dplyr) # 先把你的示例数据整合成一个数据框(设置随机种子让结果可重复) set.seed(123) df <- tibble( country = c(1,1,1,1,2,2,2,2,3,3,3,3), year = c(2000,2000,2004,2004, 2002,2002,2004,2008,2000,2000,2000,2000), party_A_number = c(11,11,12,12,21,21,22,23,31,31,31,31), party_B_number = c(12, 12, 11, 11, 22,22,21,22,32,32,32,32), party_C_number = c(13,13,13,13,23,23,23,21,33,33,33,33), party_voted = c(12,13,12,11,21,24,23,22,31,32,33,31), ideology_party_A = floor(runif(12, min=1, max=10)), ideology_party_B = floor(runif(12, min=1, max=10)), ideology_party_C = floor(runif(12, min=1, max=10)) ) # 用case_when实现条件匹配 df <- df %>% mutate(ideology_voted = case_when( party_A_number == party_voted ~ ideology_party_A, party_B_number == party_voted ~ ideology_party_B, party_C_number == party_voted ~ ideology_party_C, TRUE ~ NA_real_ # 处理没有匹配到的投票(比如示例中的24) ))
This works for 3 parties, but writing 9 lines of case_when() would be tedious. Let's move to a scalable solution.
2. 批量处理A-I政党的高效方法
For 9 parties, reshaping your data into a "long" format (instead of wide) is cleaner and easier to maintain. We'll use tidyr::pivot_longer() to restructure the party number and ideology columns, then match them to party_voted:
library(tidyr) # 定义所有政党字母:A到I party_letters <- LETTERS[1:9] # 把宽格式数据转成长格式,匹配政党编号和意识形态 df_long <- df %>% # 提取所有政党编号和意识形态列,拆分列名得到政党字母 pivot_longer( cols = starts_with("party_") & ends_with("_number") | starts_with("ideology_party_"), names_to = c(".value", "party_letter"), names_pattern = "(party|ideology_party)_(.)_(number)?" ) %>% # 修正列名,让逻辑更清晰 rename( party_number = party, ideology = ideology_party ) %>% # 只保留和投票政党匹配的行 filter(party_number == party_voted) %>% # 提取需要的列,准备合并回原始数据 select(country, year, party_voted, ideology_voted = ideology) %>% # 合并回原始数据,保留所有行(包括没有匹配的) right_join(df, by = c("country", "year", "party_voted")) %>% # 按国家和年份排序,保持数据整洁 arrange(country, year)
另一种选择:基础R循环
If you prefer working with base R instead of the tidyverse, a simple loop will also get the job done without repetitive code:
# 初始化ideology_voted列为NA df$ideology_voted <- NA_real_ # 遍历每个政党字母 for (letter in party_letters) { # 动态构造列名 number_col <- paste0("party_", letter, "_number") ideology_col <- paste0("ideology_party_", letter) # 找到匹配的行并赋值意识形态 match_indices <- df[[number_col]] == df$party_voted df$ideology_voted[match_indices] <- df[[ideology_col]][match_indices] }
关键注意事项
- 处理无匹配的情况: Both methods set
ideology_votedtoNAfor votes that don't match any party number (like24in your sample). You can adjust this by replacingNA_real_with a default value (e.g.,0) if needed. - 跨国家/年份的独立性: Since your party number-to-letter mappings are country- and year-specific, both solutions preserve this context because we're working with the full row-level data.
内容的提问来源于stack exchange,提问作者Guilherme Pires Arbache

