R语言:如何统计数据框中指定字符串的出现频率
Hey there! I see you're trying to count how often specific substrings appear in your data frame's strings, and you've already started with grepl and table—great start! Let's put those tools together to get the exact result you want.
First, let's recap your data so we're on the same page:
strings <- c("pi","pie","piece","pin","pinned","post") df <- as.data.frame(strings) counts <- c("pi", "in", "pie", "ie")
The key here is to check each target substring against every string in your data frame, then count how many times each target is found. Here's a simple, straightforward way to do it:
Step-by-Step Solution
- Use
sapplyto loop through each target in yourcountsvector. - For each target, use
grepl()to identify which strings in the data frame contain it (this returns a logical vector ofTRUE/FALSEvalues). - Sum those
TRUEvalues to get the total frequency for that target. - Combine the targets and their frequencies into a clean data frame matching your desired output.
Here's the code to make this happen:
# Calculate frequencies for each target substring frequency_results <- sapply(counts, function(target) { sum(grepl(target, df$strings)) }) # Convert the results into a properly formatted data frame final_df <- data.frame( string = names(frequency_results), freq = frequency_results, row.names = NULL # Remove default row names for clarity ) # View the final output print(final_df)
When you run this, you'll get exactly the result you're looking for:
string freq 1 pi 5 2 in 2 3 pie 2 4 ie 2
Quick Explanation
grepl(target, df$strings): Checks each string in your data frame to see if it contains the target substring. ReturnsTRUEif it does,FALSEotherwise.sum(...): Counts the number ofTRUEvalues (sinceTRUEis treated as 1 andFALSEas 0 in numeric calculations).sapply: Runs this check for every target in yourcountslist, creating a named vector of frequencies.- The final step converts that vector into a tidy data frame with the column names you specified.
If you ever need to make the search case-insensitive, just add ignore.case = TRUE to the grepl call: grepl(target, df$strings, ignore.case = TRUE).
Content of this question is from Stack Exchange, asked by JDraper

