You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言中处理电影数据集genre列的多分隔符流派拆分问题

电影数据集流派列清洗与拆分方案

问题背景

现有包含5000条数据的电影数据集,genre列存在以下问题:

  • 流派分隔符混用逗号+空格、纯空格
  • "SciFi"与"Sci-Fi"两种写法并存
  • 使用separate函数时会误将"Sci-Fi"拆分为"Sci"和"Fi"

示例数据集:

title<-c("Interstellar", "Back to the Future", "2001: A Space Odyssey", "The Martian")
genre<-c("Adventure, Drama, SciFi ", "Adventure Comedy SciFi", "Adventure, Sci-Fi", "Adventure Drama Sci-Fi")

movies<-data.frame(title, genre)

解决方案步骤

1. 统一流派名称(去除"Sci-Fi"的连字符)

用gsub批量替换所有"Sci-Fi"为"SciFi",避免拆分时误切割:

# 替换Sci-Fi为SciFi
movies$genre <- gsub("Sci-Fi", "SciFi", movies$genre)

2. 统一流派分隔符为逗号

将所有空格、逗号+空格的分隔形式统一为逗号,同时清理首尾多余空格:

# 替换所有空格或逗号+空格为逗号
movies$genre <- gsub("(, | )", ",", movies$genre)
# 去除首尾空格
movies$genre <- trimws(movies$genre)

处理后genre列内容变为:

[1] "Adventure,Drama,SciFi" "Adventure,Comedy,SciFi" "Adventure,SciFi"        "Adventure,Drama,SciFi"

3. 拆分流派为独立列

使用tidyr::separate函数拆分,注意into参数需传入字符串形式的列名:

library(tidyr)

# 拆分流派到3个独立列,不足的位置填充NA
movie_genres <- separate(movies, genre, into = c("genre1", "genre2", "genre3"), sep = ",", fill = "right")

最终结果:

title    genre1 genre2 genre3
1        Interstellar Adventure  Drama  SciFi
2 Back to the Future Adventure Comedy  SciFi
3 2001: A Space Odyssey Adventure   <NA>  SciFi
4          The Martian Adventure  Drama  SciFi

备选方案:直接按多种分隔符拆分(无需统一分隔符)

如果不想统一分隔符,可指定separate的sep参数为正则表达式,匹配逗号+空格或空格,同时先处理"Sci-Fi":

movies$genre <- gsub("Sci-Fi", "SciFi", movies$genre)
movie_genres <- separate(movies, genre, into = c("genre1", "genre2", "genre3"), sep = "(, | )", fill = "right")

内容的提问来源于stack exchange,提问作者Shashivydyula

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 02:40:29