在R语言中拆分数据框字符串并生成新行的实现方案
Solution for Splitting Long Sentences in R with dplyr
Hey Ashley, let's work through this sentence-splitting challenge together. I'll focus on a dplyr/tidyverse-based approach that follows your priority rules, updates sentence IDs correctly, and scales for larger datasets.
Step 1: Define Core Functions
First, let's set up a word-count function (you can replace this with your existing one if needed) and a recursive splitting function that applies your priority rules:
library(dplyr) library(stringr) library(tidyr) library(purrr) # Word count function count_words <- function(text) { str_count(text, "\\w+") } # Recursive function to split sentences per priority rules split_sentence_recursive <- function(sentence, threshold = 15) { # Rule 1: Split on periods (with trailing space or end of string) splits <- str_split(sentence, "\\.\\s+|\\.$")[[1]] splits <- splits[splits != ""] %>% str_trim() # Check if all splits are under the threshold word_counts <- map_int(splits, count_words) if (all(word_counts <= threshold)) { # Re-add ending periods where needed return(str_replace(splits, "$", ifelse(str_detect(sentence, "\\.$"), ".", ""))) } # Rule 2: Split on ", [Uppercase Letter]" for over-threshold splits splits <- map(splits, function(s) { if (count_words(s) > threshold) { # Split and preserve the uppercase letter after the comma split_matches <- str_extract_all(s, ",\\s+([A-Z])")[[1]] sub_splits <- str_split(s, ",\\s+([A-Z])")[[1]] %>% .[. != ""] sub_splits[-1] <- str_c(sub_splits[-1], str_remove(split_matches, ",\\s+")) # Recursively process these sub-splits split_sentence_recursive(str_c(sub_splits, collapse = " "), threshold) } else { s } }) %>% unlist() # Rule 3: Split on "and [Uppercase Letter]" if still over threshold word_counts <- map_int(splits, count_words) if (all(word_counts <= threshold)) return(splits) splits <- map(splits, function(s) { if (count_words(s) > threshold) { split_matches <- str_extract_all(s, "and\\s+([A-Z])")[[1]] sub_splits <- str_split(s, "and\\s+([A-Z])")[[1]] %>% .[. != ""] sub_splits[-1] <- str_c("and ", sub_splits[-1], str_remove(split_matches, "and\\s+")) split_sentence_recursive(str_c(sub_splits, collapse = " "), threshold) } else { s } }) %>% unlist() # Rule 4: Split on exclamation points (matches your example's "omg!") word_counts <- map_int(splits, count_words) if (all(word_counts <= threshold)) return(splits) splits <- map(splits, function(s) { if (str_detect(s, "!") && count_words(s) > threshold) { str_split(s, "!\\s+|!$")[[1]] %>% str_trim() %>% str_replace("$", "!") } else { s } }) %>% unlist() # Remove any empty strings from final splits splits[splits != ""] %>% str_trim() }
Step 2: Process the Dataset
Now we'll apply this function to your dataset, expand the split sentences into new rows, and update the sentence_id sequence:
# Process the original data frame processed_posts <- posts_sentences %>% # Group by element_id to keep sentences tied to their parent element group_by(element_id) %>% # Apply the recursive split to each sentence mutate(split_sentences = map(sentence, split_sentence_recursive, threshold = 15)) %>% # Expand split sentences into individual rows unnest(split_sentences) %>% # Recalculate word counts for each split sentence mutate(sentence_wc = count_words(split_sentences)) %>% # Reset sentence_id to be sequential per element_id group_by(element_id) %>% mutate(sentence_id = row_number()) %>% # Clean up column names and order to match your expected output rename(sentence = split_sentences) %>% select(element_id, sentence_id, sentence, sentence_wc) %>% ungroup()
Step 3: Verify the Output
You can check if the processed data matches your expected output with:
all.equal(processed_posts, expected_output)
Key Notes
- Priority Rules: The function strictly follows your specified order (periods → comma + uppercase → "and" + uppercase → exclamation points) and only moves to the next rule if the current split still produces over-threshold sentences.
- Sentence ID Handling: By grouping on
element_idand usingrow_number(), we ensure sentence IDs are sequential and correctly updated across new rows. - Scalability: This approach works for large datasets—
purrr::mapandtidyr::unnestare optimized for tidy data operations. - Customization: You can easily add more split rules (like semicolons or "but" + uppercase) by extending the recursive function with additional splitting logic.
内容的提问来源于stack exchange,提问作者Ashley
相关产品推荐
相关产品推荐

