R语言单列文本解析:移除variable列的psi#.保留剩余内容
Hey there! Let's get that variable column cleaned up properly—your regex issues are totally fixable once you know the right pattern to use.
First, let's break down what went wrong with your earlier attempts:
- The
separatecall with"psi*"uses an incorrect regex: the*here matches 0 or more of the preceding character (i), so it's only capturing "psi" (or even just "ps" or "p") instead of the full "psi[number]." prefix. - Your
str_split_fixedsyntax was off too—you mixed up the pattern and extra arguments, which is why it didn't work.
The Simple Fix with Stringr (Tidyverse)
Since you're already using dplyr, the easiest way is to use str_remove from the stringr package to strip off that unwanted prefix. The regex we need is ^psi\\d+\\.—let's break that down:
^: Matches the start of the string (so we only remove the prefix at the beginning)psi: Literally matches the characters "psi"\\d+: Matches one or more digits (the # part you mentioned)\\.: Matches the dot (we need to escape it with a backslash because dots are special in regex)
Here's the full code:
library(dplyr) library(stringr) # Clean up the variable column psi2 <- psi2 %>% mutate(variable = str_remove(variable, "^psi\\d+\\."))
Base R Alternative (No Tidyverse Needed)
If you prefer sticking to base R, use the sub function—it works the same way with the same regex:
psi2$variable <- sub("^psi\\d+\\.", "", psi2$variable)
Example to Test
Let's say your variable column has values like:c("psi1.temperature", "psi7.pressure", "psi15.wind_speed")
After running either code above, you'll get:c("temperature", "pressure", "wind_speed")
This will handle any number of digits after "psi" (whether it's 1, 2, or more) and only removes the exact prefix you want, leaving the rest of the variable name intact—perfect for when the column content changes over time!
内容的提问来源于stack exchange,提问作者Jordan

