在R中计算游戏竞赛数据集的条件round_sound均值并添加新列
游戏竞赛数据集处理需求与问题
数据集概况
拥有一个64行153列的游戏竞赛数据集,每行代表一对竞争者中的一人,相同id表示两人属同一配对组。前10行结构如下:
structure(list(id = c(20230420, 20230420, 2023042110, 2023042110, 2023042112, 2023042112, 2023042114, 2023042114, 2023042214, 2023042214), condition = c("control", "control", "control", "control", "forced_break", "forced_break", "forced_break", "forced_break", "control", "control"), round1_win = c(1, 0, 1, 0, 0, 1, 1, 0, 0, 1), round1_sound = c(1, 1, 1, 1, 3, 3, 4, 4, 2, 2), ttrs1 = c(34.8679761886597, 34.8679761886597, 23.8744249343872, 23.8744249343872, 14.2690608501434, 14.2690608501434, 17.2876904010773, 17.2876904010773, 17.2002062797546, 17.2002062797546), ttbp1 = c(42.8691244125366, 42.8691244125366, 27.8899409770966, 27.8899409770966, 17.2830331325531, 17.2830331325531, 16.2952466011047, 16.2952466011047, 21.2042264938354, 21.2042264938354), ttbi1 = c(50.7323212623596, 50.7323212623596, 34.1096398830414, 34.1096398830414, 38.0370643138885, 38.0370643138885, 36.0967524051666, 36.0967524051666, 29.5176334381103, 29.5176334381103), round2_win = c(0, 1, 1, 0, 0, 1, 0, 1, 1, 0), round2_sound = c(1, 1, 1, 1, 1, 1, 1, 1, 3, 3), ttrs2 = c(53.7323212623596, 53.7323212623596, 37.1107151508331, 37.1107151508331, 38.0380613803863, 38.0380613803863, 36.0967524051666, 36.0967524051666, 32.5185968875885, 32.5185968875885), ttbp2 = c(57.7340142726898, 57.7340142726898, 43.1248035430908, 43.1248035430908, 40.04527759552, 40.04527759552, 40.1092526912689, 40.1092526912689, 34.5237216949463, 34.5237216949463), ttbi2 = c(69.872288942337, 69.872288942337, 47.0448129177094, 47.0448129177094, 59.3737871646881, 59.3737871646881, 61.5298793315888, 61.5298793315888, 45.1512970924377, 45.1512970924377), round3_win = c(1, 0, 0, 1, 1, 0, 1, 0, 0, 1), round3_sound = c(2, 2, 1, 1, 8, 8, 8, 8, 2, 2), ttrs3 = c(72.872288942337, 72.872288942337, 50.0448129177094, 50.0448129177094, 59.3737871646881, 59.3737871646881, 61.5298793315888, 61.5298793315888, 48.1512970924377, 48.1512970924377), ttbp3 = c(79.8788452148437, 79.8788452148437, 55.0583190917969, 55.0583190917969, 65.3748495578766, 65.3748495578766, 65.538156747818, 65.538156747818, 54.1600811481476, 54.1600811481476), ttbi3 = c(92.00337266922, 92.00337266922, 59.4620923995972, 59.4620923995972, 85.2011280059815, 85.2011280059815, 85.9579682350159, 85.9579682350159, 58.6821427345276, 58.6821427345276), round4_win = c(1, 0, 0, 1, 1, 0, 1, 0, 1, 0), round4_sound = c(1, 1, 2, 2, 1, 1, 1, 1, 4, 4), ttrs4 = c(95.00337266922, 95.00337266922, 62.4620923995972, 62.4620923995972, 85.2011280059815, 85.2011280059815, 85.9579682350159, 85.9579682350159, 61.6821427345276, 61.6821427345276), ttbp4 = c(101.018859148026, 101.018859148026, 66.4772782325744, 66.4772782325744, 89.2028863430023, 89.2028863430023, 91.9701442718506, 91.9701442718506, 65.6861252784729, 65.6861252784729), ttbi4 = c(105.467720985413, 105.467720985413, 70.9975862503052, 70.9975862503052, 109.391041994095, 109.391041994095, 111.622535705566, 111.622535705566, 71.7297728061676, 71.7297728061676)), row.names = c(NA, -10L), class = c("tbl_df", "tbl", "data.frame"))
核心需求
- 计算所有
round_win==1对应的round_sound均值,添加为新列(示例命名为win_sound_mean) - 计算所有
round_win==0对应的round_sound均值,添加为另一新列(示例命名为lose_sound_mean)
当前困惑
已有相关列的编号字符串,但不知道如何运用;思路示例为rowMeans(dataset[c(3, 8, 13, 18)] **where** dataset[c(4, 9, 14, 19)] == 1),但不清楚条件筛选部分的实现。
解决方案
方法1:基础R实现
先定义对应的列索引,再逐行计算条件均值:
# 替换为实际所有round_win和round_sound列的索引 win_cols <- c(3,8,13,18) sound_cols <- c(4,9,14,19) # 计算获胜时的sound均值 dataset$win_sound_mean <- apply(dataset, 1, function(row) { win_sounds <- as.numeric(row[sound_cols])[as.numeric(row[win_cols]) == 1] if(length(win_sounds) > 0) mean(win_sounds) else NA }) # 计算失败时的sound均值 dataset$lose_sound_mean <- apply(dataset, 1, function(row) { lose_sounds <- as.numeric(row[sound_cols])[as.numeric(row[win_cols]) == 0] if(length(lose_sounds) > 0) mean(lose_sounds) else NA })
方法2:tidyverse(dplyr+purrr)实现
自动匹配列名,适合大表的管道式处理:
library(dplyr) library(purrr) # 自动提取所有round_win和round_sound列名 win_cols <- grep("^round\\d+_win$", colnames(dataset), value = TRUE) sound_cols <- grep("^round\\d+_sound$", colnames(dataset), value = TRUE) dataset <- dataset %>% rowwise() %>% mutate( win_sound_mean = mean(c_across(all_of(sound_cols))[c_across(all_of(win_cols)) == 1], na.rm = TRUE), lose_sound_mean = mean(c_across(all_of(sound_cols))[c_across(all_of(win_cols)) == 0], na.rm = TRUE) ) %>% ungroup()
内容的提问来源于stack exchange,提问作者grace.cutler
相关产品推荐
相关产品推荐

