You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Rosalind DNA基序查找问题:R代码漏检重叠匹配求助

解决Rosalind「Finding a Motif in DNA」中重叠子串匹配问题

你的问题出在gregexpr()默认采用非重叠匹配逻辑:当找到一个匹配后,会跳过已匹配的字符继续查找,导致像"ATAT"在"ATATAT"这类重叠的子串被漏检。下面提供两种可靠的修复方案:

方案1:使用正向预查正则实现重叠匹配

利用Perl兼容正则的正向零宽度预查,可以让匹配不消耗字符,从而捕捉所有重叠的起始位置。

修改后的完整代码:

# Finding a Motif in DNA

input <- readLines("input.txt", warn = FALSE) # Load data

DNAseq1 <- input[1] # 读取DNA字符串
substring1 <- input[2] # 读取待查找的子串

# 构造支持重叠匹配的正则模式:正向预查包裹目标子串
pattern <- paste0("(?=", substring1, ")")
# 启用Perl正则引擎执行匹配
matches <- gregexpr(pattern, DNAseq1, perl = TRUE)
# 提取所有匹配的起始位置
result <- unlist(matches)

print(result)

运行后针对测试用例会输出2 4 10,完全符合预期。

方案2:手动遍历字符串查找

如果你不想依赖正则,也可以通过循环逐个位置检查子串匹配,逻辑更直观:

# Finding a Motif in DNA

input <- readLines("input.txt", warn = FALSE) # Load data

DNAseq1 <- input[1]
substring1 <- input[2]

sub_len <- nchar(substring1)
dna_len <- nchar(DNAseq1)
result <- c()

# 遍历所有可能的起始位置
for (i in 1:(dna_len - sub_len + 1)) {
  # 截取当前位置开始的子串
  current_segment <- substr(DNAseq1, i, i + sub_len - 1)
  # 匹配则记录位置
  if (current_segment == substring1) {
    result <- c(result, i)
  }
}

print(result)

内容的提问来源于stack exchange,提问作者Mell

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 16:15:38