使用文件字典重编码变量类别失败,求tidyverse解决方案
问题与解决方法
问题场景
尝试使用字典app_name_recode_dictionary对数据集users_data的content_name变量进行类别重编码,使用dplyr的recode_factor函数时触发错误:
app_name必须长度为376或1,而非23
提供的字典(数据表格式)
print(app_name_recode_dictionary) content_name app_name_rename 1: DTPbrowser-prod dtp_browser_prod 2: dtp_browser_legacy_code dtp_browser_prod 3: crisprscreenViz crispr_screen_viz 4: crisprscreenviz_orig_code crispr_screen_viz 5: cage_randomiser_legacy_code cage_randomiser 6: CageRandomizer cage_randomiser 7: moprospector_legacy_code moprospector 8: moProspector moprospector 9: grow_rate_explorer_legacy_code growth_rate_explorer 10: Unknown Content growth_rate_explorer 11: ComboCor combo_cor 12: combocor_legacy_code combo_cor 13: Translatability_MultiGene translatability_multi_gene 14: multi_gene_legacy_code translatability_multi_gene 15: animals_to_groups animals_to_groups 16: animals_to_groups_legacy_code animals_to_groups 17: Translatability_Single_gene translatability_single_gene 18: single_gene_legacy_code translatability_single_gene 19: FLAT flat 20: flat_legacy_code flat 21: SyngeneicMouseBrowser syngeneic_mouse_browser 22: syngeneic_mouse_legacy_code syngeneic_mouse_browser 23: DepMapBem dep_map_bem
尝试的错误代码
factor_recode_app_name <- users_data %>% dplyr::mutate( app_name = dplyr::recode_factor( content_name, !!!app_name_recode_dictionary )) %>% dplyr::select(-content_name)
错误信息
Error in `mutate_cols()`: ! Problem with `mutate()` column `app_name`. ℹ `app_name = dplyr::recode_factor(...)`. ✖ `app_name` must be length 376 or one, not 23. Caused by error in `glubort()`: ! `app_name` must be length 376 or one, not 23.
错误原因
dplyr::recode_factor要求传入命名向量(旧类别为向量名称,新类别为对应元素),但直接传入数据表并用!!!展开后,会同时传递两列数据,导致参数长度与数据集行数不匹配(字典共23行,数据集共376行),触发报错。
Tidyverse解决方案
方案1:转换字典为命名向量后使用recode_factor
将数据表格式的字典转换为recode_factor所需的命名向量,再执行重编码:
# 生成命名向量:名称=原类别(content_name),值=新类别(app_name_rename) recode_mapping <- setNames( app_name_recode_dictionary$app_name_rename, app_name_recode_dictionary$content_name ) # 执行重编码 factor_recode_app_name <- users_data %>% dplyr::mutate(app_name = dplyr::recode_factor(content_name, !!!recode_mapping)) %>% dplyr::select(-content_name)
特点:未在字典中匹配到的content_name值会保留原始值,适合需要保留未定义类别的场景。
方案2:使用left_join实现重编码(更直观)
通过数据表连接的方式完成类别映射,未匹配的值会转为NA,便于后续识别处理:
factor_recode_app_name <- users_data %>% dplyr::left_join(app_name_recode_dictionary, by = "content_name") %>% dplyr::select(-content_name) %>% dplyr::rename(app_name = app_name_rename)
特点:逻辑清晰,易维护,适合需要明确标记未匹配类别的场景。
内容的提问来源于stack exchange,提问作者GaB
相关产品推荐
相关产品推荐

