在R中如何拆分多值电影类型vector,使每个索引仅对应单个类型?
R 实现多值电影类型向量转单值向量操作步骤
核心逻辑是先清理字符串首尾多余空格,再按分隔符拆分多值元素,最后将列表结构展开为一维向量。
基础R实现(无需加载额外包)
# 1. 模拟你提供的原始向量(实际使用时替换为你自己的向量变量) raw_genres <- c( "Action, Drama, Mystery ", "Action, Crime, Thriller ", "Action, Sci-Fi, Thriller ", "Biography, Crime, Drama ", "Action, Adventure, Drama ", "Action, Adventure, Fantasy ", "Action, Drama, Sci-Fi ", "Crime, Drama ", "Action, Sci-Fi ", "Adventure, Drama, Sci-Fi ", "Crime, Drama, Mystery ", "Action, Crime, Drama ", "Drama, Horror, Sci-Fi ", "Action, Crime, Drama ", "Comedy, Music ", "Comedy, Drama, Thriller ", "Comedy, Drama ", "Crime, Drama ", "Drama, Western ", "Crime, Drama ", "Action, Adventure, Drama ", "Action, Adventure, Thriller " ) # 2. 清理所有元素首尾的多余空格 trimmed_genres <- trimws(raw_genres) # 3. 按「逗号+任意数量空格」拆分每个多值元素,得到列表结构 split_list <- strsplit(trimmed_genres, split = ",\\s*") # 4. 展开列表得到单值向量 single_genre_vec <- unlist(split_list) # 可选:如果需要保留原始向量的索引对应关系,给结果加上原索引作为名称 names(single_genre_vec) <- rep(seq_along(raw_genres), lengths(split_list))
执行后single_genre_vec就是每个元素仅对应单个电影类型的向量,输出示例如下:
1 1 1 2 2 2 "Action" "Drama" "Mystery" "Action" "Crime" "Thriller" ...
tidyverse 实现(更适合后续数据分析联动)
library(tidyverse) single_genre_vec <- tibble(raw = raw_genres) %>% # 清理首尾空格 mutate(raw = trimws(raw)) %>% # 按分隔符拆分,每个类型单独占一行 separate_longer_delim(raw, delim = regex(",\\s*"), names_to = "genre") %>% # 提取列转为向量 pull(genre)
内容的提问来源于stack exchange,提问作者Prathamesh Damle
相关产品推荐
相关产品推荐

