在R中如何筛选数据框某列字符串以指定单词开头的行
按列筛选美国电影列表
实现指定字段筛选的代码如下:
dfUSARating <- dfUSAMovies[, c("rowNum","title", "genre", "rating", "vote")]
原有筛选逻辑的问题
你当前使用的str_detect全局包含匹配,会把所有多类型组合中带有Thriller的条目全部筛选出来,不符合你仅要首位类型为Thriller的需求:
thrill <- dfUSARating %>% filter(str_detect(dfUSARating$genre, "Thriller")) # 前3条genre字段示例输出 head(dfUSARating$genre, n=3) [1] ['Documentary', 'Comedy', 'Drama', 'Fantasy', 'Mystery', 'Sci-Fi'] [2] ['Comedy', 'Horror', 'Sci-Fi'] [3] ['Biography', 'Drama', 'Sport']
优化后的筛选方案
修改正则匹配规则,用^锚定字符串开头,仅匹配genre字段首个类型为Thriller的条目即可:
thrill_first <- dfUSARating %>% filter(str_detect(as.character(genre), "^\\['Thriller"))
说明:因为你的genre字段为factor类型,先转成character再做匹配可以避免因子level带来的匹配异常问题
你提供的样本数据结构
dput(head(dfUSARating)) structure(list(rowNum = c(6L, 7L, 8L, 12L, 13L, 15L), genre = structure(c(869L, 752L, 638L, 130L, 229L, 910L), .Label = c("['Action', 'Adventure', 'Animation', 'Comedy']", "['Action', 'Adventure', 'Biography', 'Drama', 'History', 'War']", "['Action', 'Adventure', 'Biography', 'Drama', 'History']", " ['Action', 'Adventure', 'Biography', 'History', 'Romance']", "['Action', 'Adventure', 'Biography', 'History']", "['Action', 'Adventure', 'Comedy', 'Crime', 'Drama', 'Thriller']", "['Comedy', 'Drama', 'Mystery']", "['Comedy', 'Drama', 'Romance', 'Fantasy']", "['Comedy', 'Drama', 'Romance', 'Sci-Fi']", "['Comedy', 'Drama', 'Romance', 'Sport']", "['Comedy', 'Drama', 'Romance', 'Thriller']", "['Comedy', 'Drama', 'Romance', 'War']", "['Comedy', 'Drama', 'Romance', 'Western']", "['Comedy', 'Drama', "['Western']"), class = "factor"), rating = c(5.3, 4.5, 7.8, 4.8, 7.1, 7.6)), row.names = c(NA, 6L), class = "data.frame")
内容的提问来源于stack exchange,提问作者RNewbie
相关产品推荐
相关产品推荐

