求助:如何在R中将字符串拆分为字符与数字两部分?
Got it! Splitting a string into separate character and numeric components is a common task, and R has some great tools for this—let me walk you through a couple of straightforward approaches that work for most cases.
stringr Package (Recommended for Clean Syntax) The stringr package (part of the tidyverse) makes regex-based string manipulation super intuitive. We’ll use str_extract_all to pull out all matching character/number sequences, then combine them into single strings if needed.
First, let’s define a sample vector of strings to work with:
my_strings <- c("abc123def45", "XyZ789", "987test", "a1b2c3!@#")
Extract Only Alphabetic Characters
We’ll match all sequences of uppercase/lowercase letters, then collapse any multiple matches into one string:
library(stringr) # Extract all letter sequences char_matches <- str_extract_all(my_strings, "[A-Za-z]+", simplify = TRUE) # Combine multiple matches (e.g., "abc" and "def" in the first string) characters_only <- apply(char_matches, 1, paste, collapse = "")
Extract Only Numeric Characters
Similarly, match all digit sequences and combine them:
# Extract all number sequences num_matches <- str_extract_all(my_strings, "[0-9]+", simplify = TRUE) # Combine multiple matches numbers_only <- apply(num_matches, 1, paste, collapse = "")
View the Final Result
Put it all together in a data frame for clarity:
data.frame( Original_String = my_strings, Alphabetic_Characters = characters_only, Numeric_Digits = numbers_only )
This will output:
Original_String Alphabetic_Characters Numeric_Digits 1 abc123def45 abcdef 12345 2 XyZ789 XyZ 789 3 987test test 987 4 a1b2c3!@# abc 123
If you don’t want to install additional packages, use gsub to strip out the characters you don’t want:
Extract Alphabetic Characters
Replace all digits with empty strings:
characters_only_base <- gsub("[0-9]", "", my_strings)
Extract Numeric Characters
Replace all letters with empty strings:
numbers_only_base <- gsub("[A-Za-z]", "", my_strings)
This gives the exact same result as the stringr method for our sample data!
Bonus: Handle Special Characters
If your strings include symbols (like !@# in the last example) and you want to extract all non-numeric characters (not just letters), use:
all_non_numeric <- gsub("[0-9]", "", my_strings) # For the string "a1b2c3!@#", this returns "abc!@#"
And if you want to match digits using regex shorthand, \\d works the same as [0-9]:
numbers_only_shorthand <- gsub("\\D", "", my_strings) # \\D matches non-digits
内容的提问来源于stack exchange,提问作者cephalopod

