You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中用循环或函数拆分句子为单词列及生成5元以内n元语法

Hey there! Let's work through your two R-related tasks with efficient, scalable code—no more tedious manual repetition. Here's how to tackle each one:

1. Split Sentences into a Column of Individual Words

Assuming your corpus dataframe has a column named text holding the full sentences, you can split each sentence into separate words using either base R or the tidytext package (a go-to for streamlined text processing).

Option 1: Base R Approach

# Ensure the text column is character type
corpus$text <- as.character(corpus$text)

# Split sentences into word lists and unnest into a single column
library(tidyr)
word_df <- corpus %>%
  mutate(words = strsplit(text, "\\s+")) %>%
  unnest(words)

# Preview the result (each word in its own row)
head(word_df)

Option 2: Tidytext (More Robust for Cleaning)

The tidytext package automatically handles edge cases like punctuation and lowercase conversion:

library(tidytext)
library(dplyr)

word_df <- corpus %>%
  unnest_tokens(input = text, output = word)

# This cleans up punctuation and standardizes case by default
head(word_df)
2. Generate 1-5 Grams Without Manual Work

Instead of coding each n-gram manually, we can use a loop or reusable function to generate all 1-5 grams and combine them into one dataframe. Here are two practical approaches:

Option 1: Tidytext (Simplest Method)

Tidytext's unnest_tokens has an ngram parameter that lets you define n-values. We'll loop through 1 to 5 and bind the results:

library(tidytext)
library(dplyr)
library(purrr)

# Define the range of n-grams we need
n_values <- 1:5

# Loop through each n, generate n-grams, and combine all results
all_ngrams <- map_dfr(n_values, function(n) {
  corpus %>%
    unnest_tokens(input = text, output = ngram, token = "ngrams", n = n) %>%
    mutate(n = n) # Add a column to track which n-gram type each row is
})

# Preview the combined n-grams
head(all_ngrams)

Option 2: Custom Base R Function

If you prefer avoiding tidytext, here's a custom function to generate n-grams, paired with a loop to cover 1-5 grams:

# Function to generate n-grams from a single text string
generate_ngrams <- function(text, n) {
  words <- strsplit(trimws(text), "\\s+")[[1]]
  if(length(words) < n) return(character(0)) # Skip if text is too short
  sapply(1:(length(words) - n + 1), function(i) {
    paste(words[i:(i + n - 1)], collapse = " ")
  })
}

# Apply the function to all texts for n=1 to 5
all_ngrams <- do.call(rbind, lapply(n_values, function(n) {
  grams_list <- lapply(corpus$text, generate_ngrams, n = n)
  data.frame(
    original_text = rep(corpus$text, sapply(grams_list, length)),
    ngram = unlist(grams_list),
    n = n,
    stringsAsFactors = FALSE
  )
}))

# View the final result
head(all_ngrams)

Both methods eliminate manual repetition—run the code once, and you'll have all 1-5 grams ready for analysis.

内容的提问来源于stack exchange,提问作者user6918497

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 04:08:22