如何在R中创建按年代和类型分组计数的dataframe?
按Decade和Genre分组统计IMDb电影数据的问题
需求说明
需要从IMDb筛选得到的DataFrame中,生成按decade(年代)和genre(类型)分组的计数表,预期格式如下:
| decade | genre | count |
|---|---|---|
| 1910 | Drama | 15 |
| 1920 | Drama | 27 |
| 1930 | Drama | 32 |
| ... | ... | ... |
| 1910 | Fantasy | 12 |
| 1920 | Fantasy | 23 |
| ... | ... | ... |
样本数据
structure(list(tconst = c("tt0003419", "tt0003419", "tt0004013", "tt0005231", "tt0005231", "tt0005615", "tt0005615", "tt0005772", "tt0005951", "tt0005951", "tt0006434", "tt0006434", "tt0006554", "tt0006820", "tt0007111", "tt0008826", "tt0010323", "tt0010323", "tt0010323", "tt0010323"), primaryTitle = c("The Student of Prague", "The Student of Prague", "The Ghost Breaker", "The Hound of the Baskervilles", "The Hound of the Baskervilles", "Life Without Soul", "Life Without Soul", "Mortmain", "Satan's Rhapsody", "Satan's Rhapsody", "Black Orchids", "Black Orchids", "The Crimson Stain Mystery", "Homunculus, 1. Teil", "A Night of Horror", "Alraune", "The Cabinet of Dr. Caligari", "The Cabinet of Dr. Caligari", "The Cabinet of Dr. Caligari", "The Cabinet of Dr. Caligari"), startYear = structure(c(-20819, -20819, -20454, -20089, -20089, -20089, -20089, -20089, -19358, -19358, -19358, -19358, -19724, -19724, -19358, -18628, -18263, -18263, -18263, -18263), class = "Date"), runtimeMinutes = c("85", "85", "60", "50", "50", "70", "70", "\N", "55", "55", "50", "50", "\N", "69", "56", "80", "76", "76", "76", "76"), decade = c(1910, 1910, 1910, 1910, 1910, 1910, 1910, 1910, 1910, 1910, 1910, 1910, 1910, 1910, 1910, 1910, 1920, 1920, 1920, 1920), genre = c("Drama", "Fantasy", "Adventure", "Mystery", "Crime", "Drama", "Sci-Fi", "Drama", "Fantasy", "Drama", "Drama", "Drama", "Mystery", "Sci-Fi", "Drama", "Sci-Fi", "Thriller", "Mystery", "Mystery", "Thriller" ), rating = c(6.5, 6.5, 5.2, 3.1, 3.1, 6.6, 6.6, 5.8, 6.8, 6.8, 4.8, 4.8, 6.9, 6.1, 6.1, 5.5, 8.1, 8.1, 8.1, 8.1), numVotes = c(2063, 2063, 36, 40, 40, 53, 53, 23, 719, 719, 18, 18, 18, 91, 20, 51, 62119, 62119, 62119, 62119)), row.names = c(NA, -20L), class = c("tbl_df", "tbl", "data.frame"))
尝试的代码及报错
尝试代码
subgenres %>% group_by(decade,genre) %>% summarise(count=n())
报错信息
subgenres %>% group_by(decade) %>% summarise(count=n()) Error: 'format_error' is not an exported object from 'namespace:cli' Error in count(., decade, genre) : object 'decade' not found
数据结构
tibble [6,809 × 8] (S3: tbl_df/tbl/data.frame) $ tconst : chr [1:6809] "tt0003419" "tt0003419" "tt0004013" "tt0005231" ... $ primaryTitle : chr [1:6809] "The Student of Prague" "The Student of Prague" "The Ghost Breaker" "The Hound of the Baskervilles" ... $ startYear : Date[1:6809], format: "1913-01-01" "1913-01-01" "1914-01-01" "1915-01-01" ... $ runtimeMinutes: chr [1:6809] "85" "85" "60" "50" ... $ decade : num [1:6809] 1910 1910 1910 1910 1910 1910 1910 1910 1910 1910 ... $ genre : chr [1:6809] "Drama" "Fantasy" "Adventure" "Mystery" ... $ rating : num [1:6809] 6.5 6.5 5.2 3.1 3.1 6.6 6.6 5.8 6.8 6.8 ... $ numVotes : num [1:6809] 2063 2063 36 40 40 ...
解决方案
1. 修复依赖包版本问题
报错'format_error' is not an exported object from 'namespace:cli'是因为cli包版本与tidyverse不兼容,更新相关包即可:
install.packages("tidyverse") install.packages("cli")
更新后重启R会话再执行统计代码。
2. 正确的分组统计代码
你的核心逻辑是正确的,更新包后可直接运行,也可以用更简洁的count()函数实现:
方法1:group_by + summarise
subgenres %>% group_by(decade, genre) %>% summarise(count = n(), .groups = "drop") # .groups="drop"用于取消分组状态
方法2:使用count快捷函数
subgenres %>% count(decade, genre, name = "count")
3. 可选:按类型和年代排序
如果需要按类型、年代排序,添加arrange()即可:
subgenres %>% count(decade, genre, name = "count") %>% arrange(genre, decade)
内容的提问来源于stack exchange,提问作者Tonya Howe
相关产品推荐
相关产品推荐

