You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

请求提取横杠前后文本的正则表达式及对应R代码实现

Extract Text Before/After Hyphen in R: Regex & Code Solutions

Got it, let's walk through how to extract text before and after a hyphen (-) from a character vector in R, including the regex patterns you'll need. Here's everything you need to know:

1. Regex Patterns for Hyphen Extraction

These patterns will target the text on either side of the hyphen reliably:

  • Extract text before the hyphen: ^[^-]+

    • ^: Anchors the match to the start of the string (so we don't accidentally match text later in the string if there were multiple hyphens)
    • [^-]+: Matches one or more characters that are not a hyphen (the [^...] syntax is a negated character class, which excludes the hyphen itself)
  • Extract text after the hyphen: [^-]+$

    • [^-]+: Matches one or more non-hyphen characters
    • $: Anchors the match to the end of the string (ensures we get everything after the last hyphen, though in your example each string only has one)

If you want to capture both parts in a single pass (useful for more complex parsing scenarios), use a pattern with capture groups: ^(.*?)-(.*)$

  • (.*?): Non-greedy capture of all characters before the first hyphen (the ? stops us at the first hyphen, avoiding issues if a string had multiple hyphens)
  • (.*): Captures all characters after the hyphen

2. R Code to Extract Hyphenated Text

Let's use your example vector vec <- c("1-2", "3-4", "5-6") to demonstrate three common approaches:

Method 1: Using stringr (Tidyverse-Friendly)

This is the most intuitive approach if you're using the tidyverse. First install/load the package:

# Install if you haven't already
# install.packages("stringr")
library(stringr)

# Extract text BEFORE the hyphen
before_hyphen <- str_extract(vec, "^[^-]+")
# Output: "1" "3" "5"

# Extract text AFTER the hyphen
after_hyphen <- str_extract(vec, "[^-]+$")
# Output: "2" "4" "6"

Method 2: Base R with strsplit

No external packages needed here—split each string at the hyphen, then pull the first/second element from each split result:

# Extract before hyphen
before_base <- sapply(strsplit(vec, "-"), `[`, 1)

# Extract after hyphen
after_base <- sapply(strsplit(vec, "-"), `[`, 2)

Method 3: Base R with Regex Capture Groups

If you prefer using the capture group pattern, use regexec to define the match and regmatches to extract the groups:

# Define the regex with capture groups and match against the vector
matches <- regexec("^(.*?)-(.*)$", vec)
captured_text <- regmatches(vec, matches)

# Extract the first capture group (text before hyphen)
before_capture <- sapply(captured_text, `[`, 2)

# Extract the second capture group (text after hyphen)
after_capture <- sapply(captured_text, `[`, 3)

All three methods will give you the exact results you're looking for with your example vector.

内容的提问来源于stack exchange,提问作者Andrea Ferrante

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 04:00:48