You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Base R中提取字符串中'_'字符之前的内容?

Extract Text Before First Underscore in Base R

Hey there! Let's sort out this string extraction problem for you. You're trying to grab the part of each string before the first underscore in your vector, and your current substr + gregexpr approach isn't working as expected—here's why, plus two solid Base R solutions:

Why Your Original Code Fails

The issue with substr(a,1,unlist(gregexpr(pattern ='_',a))-1) is that gregexpr returns all positions of underscores for each string. For example, your second string "NM_020519_1" has two underscores, so gregexpr returns two positions for it. When you unlist this, you end up with a longer vector than your original a, causing substr to recycle values incorrectly and produce wrong results.

Solution 1: Use sub() (Simplest & Most Efficient)

The easiest way in Base R is to use sub() with a regular expression that matches the first underscore and everything after it, then replace that match with an empty string:

a <- c("NM020506_1","NM_020519_1","NM00_1030297.2")
result <- sub("_.*", "", a)
result
# Output: [1] "NM020506" "NM"       "NM00"
  • _.* matches the first underscore (_) followed by any number of characters (.*)
  • sub() only replaces the first occurrence of this pattern, which is exactly what we need to keep just the text before the first underscore.

Solution 2: Fix Your gregexpr Approach

If you prefer working with gregexpr and substr, you need to target only the first underscore position for each string. Use sapply to process each element individually:

result <- sapply(a, function(x) {
  # Get the position of the first underscore
  first_underscore_pos <- gregexpr("_", x)[[1]][1]
  # If no underscore exists, return the original string; else extract up to before the underscore
  if (is.na(first_underscore_pos)) x else substr(x, 1, first_underscore_pos - 1)
})
# Remove the names added by sapply
unname(result)
# Output: [1] "NM020506" "NM"       "NM00"

This ensures we're only using the first underscore's position for each string, avoiding the recycling issue from your original code.

内容的提问来源于stack exchange,提问作者itthrill

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:22:17