在R中使用stringi提取字符串中特定字符后的多个子字符串
Hey there! I see you need to pull out all those terms that follow wiki/ in each sentence of your data frame. Let's use the stringi package to get this done cleanly and efficiently.
Step 1: Install and Load the stringi Package
First, make sure you have the package installed (if you haven't already) and loaded into your R session:
install.packages("stringi") library(stringi)
Step 2: Extract the Target Substrings
We'll use stri_extract_all_regex() to find all matches in each string. The regex pattern targets wiki/ followed by any sequence of non-space characters (since your terms are separated by spaces). Setting capture_groups = TRUE ensures we only get the part after wiki/, not the full matched string:
# Apply the extraction to your data frame data$extracted_terms <- stri_extract_all_regex( str = data$datalist, pattern = "wiki/([^\\s]+)", capture_groups = TRUE )
What the Result Looks Like
The extracted_terms column will be a list column where each entry is a vector of the terms from the corresponding sentence. For example:
- The first row will return
c("political_philosophy", "self-governance") - The second row will return
c("stateless_society", "hierarchy", "free_association_(communism_and_anarchism)")
Optional: Convert to Long Format
If you prefer each extracted term to be on its own row (instead of nested in a list), you can use tidyr::unnest() to expand the data frame:
# Load tidyr if you haven't already install.packages("tidyr") library(tidyr) # Expand to long format data_long <- data %>% unnest(extracted_terms)
That's it! This approach works smoothly even for large data frames, as stringi is optimized for fast string operations.
内容的提问来源于stack exchange,提问作者Kayla

