R语言:无循环提取DataFrame文本列Top2高频词并添加二进制列
Got it, let's tackle this problem step by step with clean, loop-free R code—no messy iterations required! We'll use the tidyverse toolkit (dplyr, tidyr, stringr) since it’s perfect for concise text manipulation. If you don’t have it installed, run install.packages("tidyverse") first to get set up.
Step 1: Identify the Top 2 Most Frequent Words
First, we’ll split each text entry into individual words, count how often each word appears, and grab the top 2:
library(tidyverse) # Your original data frame df <- data.frame( text = c("test and run", "rest and sleep", "test", "test of course"), id = c('a','b','c','d') ) # Extract and rank top 2 words top_words <- df %>% unnest_tokens(word, text) %>% # Split text into single words count(word, sort = TRUE) %>% # Count occurrences, sorted descending slice_head(n = 2) %>% # Keep top 2 pull(word) # Extract as a vector: c("test", "and")
Step 2: Add Binary Columns for Each Top Word
Next, we’ll create binary indicators (1 = word present, 0 = not present) for each of the top words. We can either make separate columns or combine them into a single comma-separated column like your example:
# Add individual binary columns df <- df %>% mutate( has_test = as.integer(str_detect(text, fixed(top_words[1]))), has_and = as.integer(str_detect(text, fixed(top_words[2]))) ) # Optional: Combine into a single comma-separated column df <- df %>% mutate(topTextBinary = paste(has_test, has_and, sep = ","))
Final Result
Running this code gives you exactly the output you wanted:
text id has_test has_and topTextBinary 1 test and run a 1 1 1,1 2 rest and sleep b 0 1 0,1 3 test c 1 0 1,0 4 test of course d 1 0 1,0
Pro Tip for Dynamic Scaling
If you want a more flexible solution (say, if you might need top 3 or 5 words later), use purrr::map_dfc to generate binary columns automatically without hardcoding word names:
df <- df %>% bind_cols( map_dfc(set_names(top_words), ~as.integer(str_detect(df$text, fixed(.x)))) ) %>% mutate(topTextBinary = paste(!!!syms(top_words), sep = ","))
This way, if your top words ever change, the column names and binary checks update automatically.
内容的提问来源于stack exchange,提问作者Stefania Axo

