You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中使用stringi提取字符串中特定字符后的多个子字符串

Extract Substrings After "wiki/" Using stringi in R

Hey there! I see you need to pull out all those terms that follow wiki/ in each sentence of your data frame. Let's use the stringi package to get this done cleanly and efficiently.

Step 1: Install and Load the stringi Package

First, make sure you have the package installed (if you haven't already) and loaded into your R session:

install.packages("stringi")
library(stringi)

Step 2: Extract the Target Substrings

We'll use stri_extract_all_regex() to find all matches in each string. The regex pattern targets wiki/ followed by any sequence of non-space characters (since your terms are separated by spaces). Setting capture_groups = TRUE ensures we only get the part after wiki/, not the full matched string:

# Apply the extraction to your data frame
data$extracted_terms <- stri_extract_all_regex(
  str = data$datalist,
  pattern = "wiki/([^\\s]+)",
  capture_groups = TRUE
)

What the Result Looks Like

The extracted_terms column will be a list column where each entry is a vector of the terms from the corresponding sentence. For example:

  • The first row will return c("political_philosophy", "self-governance")
  • The second row will return c("stateless_society", "hierarchy", "free_association_(communism_and_anarchism)")

Optional: Convert to Long Format

If you prefer each extracted term to be on its own row (instead of nested in a list), you can use tidyr::unnest() to expand the data frame:

# Load tidyr if you haven't already
install.packages("tidyr")
library(tidyr)

# Expand to long format
data_long <- data %>%
  unnest(extracted_terms)

That's it! This approach works smoothly even for large data frames, as stringi is optimized for fast string operations.

内容的提问来源于stack exchange,提问作者Kayla

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 04:21:08