如何使用tidyr分割DataFrame列时保留分隔符?
Hey there! Let's fix that missing opening quote issue in your text column. The problem with your original approaches is that you were either excluding the quote from your capture group or treating it as part of the separator (which gets removed). Here are two straightforward solutions that'll keep that quote intact:
Solution 1: Use extract with a corrected regex
We'll adjust the regular expression to explicitly capture the opening quote along with the text inside it. This way, the entire quoted string (including both quotes) lands in the text column.
library(tidyr) library(dplyr) # Replicate your sample data df <- tibble( ID = c(1, 2), Value = c('message "some text"', 'more messages "some more text"') ) # Extract with a regex that captures the full quoted string df_clean <- df %>% extract( col = Value, into = c("message", "text"), regex = '^(.*?) ("[^"]+")$', # Second group captures the full quoted text remove = TRUE ) print(df_clean)
Solution 2: Use separate with a lookahead regex
If you prefer separate, we can use a positive lookahead to split only at the space that comes right before an opening quote. This ensures the quote itself isn't treated as part of the separator and stays attached to the text.
df_clean <- df %>% separate( Value, into = c("message", "text"), sep = ' (?=")', # Split at space followed by ", without consuming the " remove = TRUE ) print(df_clean)
Why your original attempts didn't work:
- Your first
separatecall used' "'as the separator, which includes the opening quote—so that quote gets discarded along with the separator. - Your
extractregex^(.*?) "(.*?)$only captured the text after the opening quote, not the quote itself. The corrected regex fixes this by wrapping the quoted portion in its own capture group.
Both solutions will give you the desired output where the text column retains the full "some text" format.
内容的提问来源于stack exchange,提问作者french_fries

