You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言如何定位V1列'-'连续区间并匹配V2列对应内容

核心思路

核心是用正则表达式-+匹配连续出现的-字符,直接获取连续区间的起止位置,无需手动合并单个-的定位结果。

依赖包

需要用到stringr做字符串匹配/提取,dplyr和purrr做逐行处理:

library(dplyr)
library(stringr)
library(purrr)

处理代码

res <- ip %>%
  # 逐行处理
  rowwise() %>%
  mutate(
    # 匹配所有连续'-'的起止区间
    dash_pos = list(str_locate(V1, "-+")),
    # 逐个区间拼接要求格式的字符串
    seg_str = list(map2_chr(
      dash_pos[, "start"], dash_pos[, "end"],
      \(s, e) {
        # 提取V2对应起止位置的字符
        start_char = str_sub(V2, s, s)
        end_char = str_sub(V2, e, e)
        # 计算区间长度
        len = e - s + 1
        # 按要求拼接
        sprintf("%s:%s-%d:%d-%d", start_char, end_char, s, e, len)
      }
    )),
    # 多个区间用分号分隔
    output = paste(seg_str, collapse = "; ")
  ) %>%
  ungroup() %>%
  # 保留需要的输出列
  select(V1, V2, output)

如果使用R4.1以下版本,将代码中的\(s, e)替换为function(s, e)即可。

结果验证

输出的res和你给出的预期op完全一致。


内容的提问来源于stack exchange,提问作者sdu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 19:18:01