在R语言中从字符向量统计《弗兰肯斯坦》文本字符出现频率
Hey there! Since you've already split your Frankenstein.txt content into a character vector character_array, here are a few straightforward, efficient ways to count how often each character appears:
table() Function This is the simplest, most direct approach—base R's built-in table() function is made exactly for frequency counting. Just one line gets you started:
char_frequencies <- table(character_array)
The result will be a table object where each row pairs a character with its total occurrences. If you want to convert this into a more flexible data frame (great for sorting or further analysis), add:
char_freq_df <- as.data.frame(char_frequencies, stringsAsFactors = FALSE) colnames(char_freq_df) <- c("Character", "Count")
dplyr::count() for Tidyverse Workflows If you prefer working with the tidyverse ecosystem, you can turn your character vector into a tibble and use count() to tally occurrences—plus, you can sort results in one step:
library(dplyr) char_freq_tbl <- tibble(character = character_array) %>% count(character, sort = TRUE)
The sort = TRUE argument will automatically order results from most to least frequent, saving you an extra sorting step later.
Your sample vector includes \r (a carriage return character)—you might want to exclude these or other invisible control characters (like \n or whitespace) from your counts. Here's how to filter them first:
Base R Version
# Filter out whitespace, carriage returns, and line breaks clean_chars <- character_array[!grepl("^\\s$|\\r|\\n", character_array)] clean_counts <- table(clean_chars)
Tidyverse Version
clean_freq_tbl <- tibble(character = character_array) %>% filter(!grepl("^\\s$|\\r|\\n", character)) %>% count(character, sort = TRUE)
These methods should give you exactly the character count breakdown you need for your Frankenstein text!
内容的提问来源于stack exchange,提问作者Paul Trimor

