You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于CSES数据集批量构建ideology_voted变量的技术问询

解决CSES数据集的ideology_voted变量构建问题

Hey there! Let's work through this problem together. You need to create a variable ideology_voted that maps each respondent's voted party to its corresponding self-perceived ideology—and you want to avoid writing repetitive conditional code for all 9 parties (A-I). Let's break this down into two parts: first fixing the repeated conditional logic you're struggling with, then scaling it up for all parties.

1. 先解决重复条件的基础写法

If you're just starting out, using dplyr::case_when() makes the conditional logic clear and readable, even for a few parties. Here's how to apply it to your sample data:

library(dplyr)

# 先把你的示例数据整合成一个数据框(设置随机种子让结果可重复)
set.seed(123)
df <- tibble(
  country = c(1,1,1,1,2,2,2,2,3,3,3,3),
  year = c(2000,2000,2004,2004, 2002,2002,2004,2008,2000,2000,2000,2000),
  party_A_number = c(11,11,12,12,21,21,22,23,31,31,31,31),
  party_B_number = c(12, 12, 11, 11, 22,22,21,22,32,32,32,32),
  party_C_number = c(13,13,13,13,23,23,23,21,33,33,33,33),
  party_voted = c(12,13,12,11,21,24,23,22,31,32,33,31),
  ideology_party_A = floor(runif(12, min=1, max=10)),
  ideology_party_B = floor(runif(12, min=1, max=10)),
  ideology_party_C = floor(runif(12, min=1, max=10))
)

# 用case_when实现条件匹配
df <- df %>%
  mutate(ideology_voted = case_when(
    party_A_number == party_voted ~ ideology_party_A,
    party_B_number == party_voted ~ ideology_party_B,
    party_C_number == party_voted ~ ideology_party_C,
    TRUE ~ NA_real_ # 处理没有匹配到的投票(比如示例中的24)
  ))

This works for 3 parties, but writing 9 lines of case_when() would be tedious. Let's move to a scalable solution.

2. 批量处理A-I政党的高效方法

For 9 parties, reshaping your data into a "long" format (instead of wide) is cleaner and easier to maintain. We'll use tidyr::pivot_longer() to restructure the party number and ideology columns, then match them to party_voted:

library(tidyr)

# 定义所有政党字母:A到I
party_letters <- LETTERS[1:9]

# 把宽格式数据转成长格式,匹配政党编号和意识形态
df_long <- df %>%
  # 提取所有政党编号和意识形态列,拆分列名得到政党字母
  pivot_longer(
    cols = starts_with("party_") & ends_with("_number") | starts_with("ideology_party_"),
    names_to = c(".value", "party_letter"),
    names_pattern = "(party|ideology_party)_(.)_(number)?"
  ) %>%
  # 修正列名,让逻辑更清晰
  rename(
    party_number = party,
    ideology = ideology_party
  ) %>%
  # 只保留和投票政党匹配的行
  filter(party_number == party_voted) %>%
  # 提取需要的列,准备合并回原始数据
  select(country, year, party_voted, ideology_voted = ideology) %>%
  # 合并回原始数据,保留所有行(包括没有匹配的)
  right_join(df, by = c("country", "year", "party_voted")) %>%
  # 按国家和年份排序,保持数据整洁
  arrange(country, year)

另一种选择:基础R循环

If you prefer working with base R instead of the tidyverse, a simple loop will also get the job done without repetitive code:

# 初始化ideology_voted列为NA
df$ideology_voted <- NA_real_

# 遍历每个政党字母
for (letter in party_letters) {
  # 动态构造列名
  number_col <- paste0("party_", letter, "_number")
  ideology_col <- paste0("ideology_party_", letter)
  
  # 找到匹配的行并赋值意识形态
  match_indices <- df[[number_col]] == df$party_voted
  df$ideology_voted[match_indices] <- df[[ideology_col]][match_indices]
}

关键注意事项

  • 处理无匹配的情况: Both methods set ideology_voted to NA for votes that don't match any party number (like 24 in your sample). You can adjust this by replacing NA_real_ with a default value (e.g., 0) if needed.
  • 跨国家/年份的独立性: Since your party number-to-letter mappings are country- and year-specific, both solutions preserve this context because we're working with the full row-level data.

内容的提问来源于stack exchange,提问作者Guilherme Pires Arbache

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 08:48:55