R语言技术问询:如何提取字符串首个句号/冒号前的内容?
Got it! Let's tackle this problem—you need to pull the text before the first . or : in each element of your R vector. Here are two straightforward approaches:
1. Base R 原生方法(无需额外安装包)
Use the built-in sub() function with a regex pattern to strip away everything from the first separator onward:
city <- c('Kirkland-1234.It is a goodtown','Bethesda-345. small town', 'Wellington: 12345') # Extract text before first . or : result_base <- sub("([^.:]+).*", "\\1", city) print(result_base) # Output: "Kirkland-1234" "Bethesda-345" "Wellington"
Regex breakdown:
([^.:]+): This is a capture group that matches one or more characters that are NOT.or:(the[^...]syntax creates an exclusion character class). It stops as soon as it hits the first.or:..*: Matches everything after the separator.\\1: Replaces the entire string with the content from the first capture group—exactly the text we want.
2. stringr Package Method (more readable)
If you prefer the stringr toolkit for string manipulation, str_extract() makes this even simpler by directly grabbing the matching portion:
library(stringr) result_str <- str_extract(city, "^[^.:]+") print(result_str) # Same output as the base R method
Regex breakdown:
^: Anchors the match to the start of the string, ensuring we don't accidentally match separators later in the text.[^.:]+: Again, matches all characters up to the first.or:.
Bonus note:
If any element in your vector doesn't have a . or :, both methods will return the entire string—which is exactly what you'd want in that case.
内容的提问来源于stack exchange,提问作者user3570187
相关产品推荐
相关产品推荐

