在R中用循环或函数拆分句子为单词列及生成5元以内n元语法
Hey there! Let's work through your two R-related tasks with efficient, scalable code—no more tedious manual repetition. Here's how to tackle each one:
Assuming your corpus dataframe has a column named text holding the full sentences, you can split each sentence into separate words using either base R or the tidytext package (a go-to for streamlined text processing).
Option 1: Base R Approach
# Ensure the text column is character type corpus$text <- as.character(corpus$text) # Split sentences into word lists and unnest into a single column library(tidyr) word_df <- corpus %>% mutate(words = strsplit(text, "\\s+")) %>% unnest(words) # Preview the result (each word in its own row) head(word_df)
Option 2: Tidytext (More Robust for Cleaning)
The tidytext package automatically handles edge cases like punctuation and lowercase conversion:
library(tidytext) library(dplyr) word_df <- corpus %>% unnest_tokens(input = text, output = word) # This cleans up punctuation and standardizes case by default head(word_df)
Instead of coding each n-gram manually, we can use a loop or reusable function to generate all 1-5 grams and combine them into one dataframe. Here are two practical approaches:
Option 1: Tidytext (Simplest Method)
Tidytext's unnest_tokens has an ngram parameter that lets you define n-values. We'll loop through 1 to 5 and bind the results:
library(tidytext) library(dplyr) library(purrr) # Define the range of n-grams we need n_values <- 1:5 # Loop through each n, generate n-grams, and combine all results all_ngrams <- map_dfr(n_values, function(n) { corpus %>% unnest_tokens(input = text, output = ngram, token = "ngrams", n = n) %>% mutate(n = n) # Add a column to track which n-gram type each row is }) # Preview the combined n-grams head(all_ngrams)
Option 2: Custom Base R Function
If you prefer avoiding tidytext, here's a custom function to generate n-grams, paired with a loop to cover 1-5 grams:
# Function to generate n-grams from a single text string generate_ngrams <- function(text, n) { words <- strsplit(trimws(text), "\\s+")[[1]] if(length(words) < n) return(character(0)) # Skip if text is too short sapply(1:(length(words) - n + 1), function(i) { paste(words[i:(i + n - 1)], collapse = " ") }) } # Apply the function to all texts for n=1 to 5 all_ngrams <- do.call(rbind, lapply(n_values, function(n) { grams_list <- lapply(corpus$text, generate_ngrams, n = n) data.frame( original_text = rep(corpus$text, sapply(grams_list, length)), ngram = unlist(grams_list), n = n, stringsAsFactors = FALSE ) })) # View the final result head(all_ngrams)
Both methods eliminate manual repetition—run the code once, and you'll have all 1-5 grams ready for analysis.
内容的提问来源于stack exchange,提问作者user6918497

