基于多列数据生成新列:皮肤问题位置数据集重构需求
重构皮肤问题位置数据集:多列转部位二进制列
嘿,我来帮你搞定这个数据格式转换的需求!看起来你需要把分散在20个scaloc列里的皮肤问题位置,转换成按身体部位分类的二进制列(1表示该部位有问题,0表示没有)。我给你两种常用工具的实现方法,你可以根据自己用的语言选择:
方法1:用R(dplyr/tidyr)
步骤说明
- 先定义位置编号到身体部位的映射(你可以根据实际对应关系修改)
- 要么通过「宽转长再转宽」的方式处理,要么直接用
rowSums快速生成二进制列
快速实现(rowSums法)
这种方法不需要转换数据格式,直接逐行检查是否存在目标位置:
library(dplyr) # 假设你的数据集名为skin_data # 第一步:定义位置-部位映射 location_to_part <- list( face = c(3), # 示例:位置3对应face torso = c(5, 6, 7), # 示例:位置5/6/7对应torso arms = c(10, 11), # 示例:位置10/11对应arms legs = c(15, 16) # 示例:位置15/16对应legs ) # 第二步:生成各部位的二进制列 skin_data_processed <- skin_data %>% mutate( face = as.integer(rowSums(select(., starts_with("scaloc")) == location_to_part$face, na.rm = TRUE) > 0), torso = as.integer(rowSums(select(., starts_with("scaloc")) %in% location_to_part$torso, na.rm = TRUE) > 0), arms = as.integer(rowSums(select(., starts_with("scaloc")) %in% location_to_part$arms, na.rm = TRUE) > 0), legs = as.integer(rowSums(select(., starts_with("scaloc")) %in% location_to_part$legs, na.rm = TRUE) > 0) )
灵活扩展(宽转长法)
如果后续需要调整部位分类,这种方法更易维护:
library(dplyr) library(tidyr) skin_data_processed <- skin_data %>% # 把20个scaloc列转成长格式,丢弃NA值 pivot_longer( cols = starts_with("scaloc"), names_to = "col_name", values_to = "location", values_drop_na = TRUE ) %>% # 给每个位置匹配对应的身体部位 mutate(body_part = case_when( location %in% c(3) ~ "face", location %in% c(5,6,7) ~ "torso", location %in% c(10,11) ~ "arms", location %in% c(15,16) ~ "legs", TRUE ~ NA_character_ # 未定义的位置标记为NA )) %>% # 按原始行分组,每个部位标记为1 group_by(across(-c(col_name, location, body_part))) %>% summarise( face = max(as.integer(body_part == "face"), na.rm = TRUE), torso = max(as.integer(body_part == "torso"), na.rm = TRUE), arms = max(as.integer(body_part == "arms"), na.rm = TRUE), legs = max(as.integer(body_part == "legs"), na.rm = TRUE), .groups = "drop" ) %>% # 和原始数据合并,确保未出现部位的行设为0 right_join(skin_data, by = names(skin_data)[!grepl("scaloc", names(skin_data))]) %>% mutate(across(c(face, torso, arms, legs), ~replace_na(., 0)))
方法2:用Python(Pandas)
Pandas的isin+any方法可以快速实现需求,代码简洁易读:
import pandas as pd # 假设你的数据集名为skin_df # 定义位置-部位映射 location_mapping = { 'face': [3], 'torso': [5, 6, 7], 'arms': [10, 11], 'legs': [15, 16] } # 提取所有scaloc列 scaloc_columns = [col for col in skin_df.columns if col.startswith('scaloc')] # 循环生成每个部位的二进制列 for part, locations in location_mapping.items(): # 检查每行是否有属于该部位的位置,转成整数1/0 skin_df[part] = skin_df[scaloc_columns].isin(locations).any(axis=1).astype(int)
关键提示
- 你只需要根据实际业务逻辑,修改位置-部位的映射列表即可,比如如果face对应位置3、4,就把
face = [3]改成face = [3,4] - 两种方法都会自动处理
NA值,只要某一行的scaloc列中有任意一个属于目标部位的值,就会将对应部位列设为1
内容的提问来源于stack exchange,提问作者mmarks
相关产品推荐
相关产品推荐

