You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R语言中拆分数据框字符串并生成新行的实现方案

Solution for Splitting Long Sentences in R with dplyr

Hey Ashley, let's work through this sentence-splitting challenge together. I'll focus on a dplyr/tidyverse-based approach that follows your priority rules, updates sentence IDs correctly, and scales for larger datasets.

Step 1: Define Core Functions

First, let's set up a word-count function (you can replace this with your existing one if needed) and a recursive splitting function that applies your priority rules:

library(dplyr)
library(stringr)
library(tidyr)
library(purrr)

# Word count function
count_words <- function(text) {
  str_count(text, "\\w+")
}

# Recursive function to split sentences per priority rules
split_sentence_recursive <- function(sentence, threshold = 15) {
  # Rule 1: Split on periods (with trailing space or end of string)
  splits <- str_split(sentence, "\\.\\s+|\\.$")[[1]]
  splits <- splits[splits != ""] %>% str_trim()
  
  # Check if all splits are under the threshold
  word_counts <- map_int(splits, count_words)
  if (all(word_counts <= threshold)) {
    # Re-add ending periods where needed
    return(str_replace(splits, "$", ifelse(str_detect(sentence, "\\.$"), ".", "")))
  }
  
  # Rule 2: Split on ", [Uppercase Letter]" for over-threshold splits
  splits <- map(splits, function(s) {
    if (count_words(s) > threshold) {
      # Split and preserve the uppercase letter after the comma
      split_matches <- str_extract_all(s, ",\\s+([A-Z])")[[1]]
      sub_splits <- str_split(s, ",\\s+([A-Z])")[[1]] %>% .[. != ""]
      sub_splits[-1] <- str_c(sub_splits[-1], str_remove(split_matches, ",\\s+"))
      # Recursively process these sub-splits
      split_sentence_recursive(str_c(sub_splits, collapse = " "), threshold)
    } else {
      s
    }
  }) %>% unlist()
  
  # Rule 3: Split on "and [Uppercase Letter]" if still over threshold
  word_counts <- map_int(splits, count_words)
  if (all(word_counts <= threshold)) return(splits)
  
  splits <- map(splits, function(s) {
    if (count_words(s) > threshold) {
      split_matches <- str_extract_all(s, "and\\s+([A-Z])")[[1]]
      sub_splits <- str_split(s, "and\\s+([A-Z])")[[1]] %>% .[. != ""]
      sub_splits[-1] <- str_c("and ", sub_splits[-1], str_remove(split_matches, "and\\s+"))
      split_sentence_recursive(str_c(sub_splits, collapse = " "), threshold)
    } else {
      s
    }
  }) %>% unlist()
  
  # Rule 4: Split on exclamation points (matches your example's "omg!")
  word_counts <- map_int(splits, count_words)
  if (all(word_counts <= threshold)) return(splits)
  
  splits <- map(splits, function(s) {
    if (str_detect(s, "!") && count_words(s) > threshold) {
      str_split(s, "!\\s+|!$")[[1]] %>% str_trim() %>% str_replace("$", "!")
    } else {
      s
    }
  }) %>% unlist()
  
  # Remove any empty strings from final splits
  splits[splits != ""] %>% str_trim()
}

Step 2: Process the Dataset

Now we'll apply this function to your dataset, expand the split sentences into new rows, and update the sentence_id sequence:

# Process the original data frame
processed_posts <- posts_sentences %>%
  # Group by element_id to keep sentences tied to their parent element
  group_by(element_id) %>%
  # Apply the recursive split to each sentence
  mutate(split_sentences = map(sentence, split_sentence_recursive, threshold = 15)) %>%
  # Expand split sentences into individual rows
  unnest(split_sentences) %>%
  # Recalculate word counts for each split sentence
  mutate(sentence_wc = count_words(split_sentences)) %>%
  # Reset sentence_id to be sequential per element_id
  group_by(element_id) %>%
  mutate(sentence_id = row_number()) %>%
  # Clean up column names and order to match your expected output
  rename(sentence = split_sentences) %>%
  select(element_id, sentence_id, sentence, sentence_wc) %>%
  ungroup()

Step 3: Verify the Output

You can check if the processed data matches your expected output with:

all.equal(processed_posts, expected_output)

Key Notes

  • Priority Rules: The function strictly follows your specified order (periods → comma + uppercase → "and" + uppercase → exclamation points) and only moves to the next rule if the current split still produces over-threshold sentences.
  • Sentence ID Handling: By grouping on element_id and using row_number(), we ensure sentence IDs are sequential and correctly updated across new rows.
  • Scalability: This approach works for large datasets—purrr::map and tidyr::unnest are optimized for tidy data operations.
  • Customization: You can easily add more split rules (like semicolons or "but" + uppercase) by extending the recursive function with additional splitting logic.

内容的提问来源于stack exchange,提问作者Ashley

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 07:08:36