You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言:无循环提取DataFrame文本列Top2高频词并添加二进制列

Solution for Extracting Top 2 Words & Adding Binary Columns

Got it, let's tackle this problem step by step with clean, loop-free R code—no messy iterations required! We'll use the tidyverse toolkit (dplyr, tidyr, stringr) since it’s perfect for concise text manipulation. If you don’t have it installed, run install.packages("tidyverse") first to get set up.

Step 1: Identify the Top 2 Most Frequent Words

First, we’ll split each text entry into individual words, count how often each word appears, and grab the top 2:

library(tidyverse)

# Your original data frame
df <- data.frame(
  text = c("test and run", "rest and sleep", "test", "test of course"),
  id = c('a','b','c','d')
)

# Extract and rank top 2 words
top_words <- df %>%
  unnest_tokens(word, text) %>%  # Split text into single words
  count(word, sort = TRUE) %>%   # Count occurrences, sorted descending
  slice_head(n = 2) %>%          # Keep top 2
  pull(word)                     # Extract as a vector: c("test", "and")

Step 2: Add Binary Columns for Each Top Word

Next, we’ll create binary indicators (1 = word present, 0 = not present) for each of the top words. We can either make separate columns or combine them into a single comma-separated column like your example:

# Add individual binary columns
df <- df %>%
  mutate(
    has_test = as.integer(str_detect(text, fixed(top_words[1]))),
    has_and = as.integer(str_detect(text, fixed(top_words[2])))
  )

# Optional: Combine into a single comma-separated column
df <- df %>%
  mutate(topTextBinary = paste(has_test, has_and, sep = ","))

Final Result

Running this code gives you exactly the output you wanted:

text id has_test has_and topTextBinary
1   test and run  a        1       1           1,1
2 rest and sleep  b        0       1           0,1
3           test  c        1       0           1,0
4 test of course  d        1       0           1,0

Pro Tip for Dynamic Scaling

If you want a more flexible solution (say, if you might need top 3 or 5 words later), use purrr::map_dfc to generate binary columns automatically without hardcoding word names:

df <- df %>%
  bind_cols(
    map_dfc(set_names(top_words), ~as.integer(str_detect(df$text, fixed(.x))))
  ) %>%
  mutate(topTextBinary = paste(!!!syms(top_words), sep = ","))

This way, if your top words ever change, the column names and binary checks update automatically.

内容的提问来源于stack exchange,提问作者Stefania Axo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:53:55