You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从AAStringSet对象中筛选指定位置的氨基酸序列?

问题描述

我有如下AAStringSet对象,希望从seq列中筛选指定区域的序列:

df <- AAStringSet(c("MEKIVLLLA", "MEKIVLDIA"))

输入的AAStringSet对象长度为2:

width  seq        names
   9    MEKIVLLLA  A1
   9    MEKIVLDIA  A2

期望输出的AAStringSet对象长度为2:

width  seq        names
   6    MKLLLA  A1
   6    MKLDIA  A2

我使用以下代码但无法正常运行:

positions <- c(1,3,6,7,8,9)
end_positions <- c(1,3,6,7,8,9)
subseq(df, start = positions, end = end_positions))
解决方案

你用subseq()的思路不对——这个函数的start和end参数是用来指定连续序列区间的,而你需要提取的是非连续的单个位置集合,这时候应该用Biostrings包中的extractAt()函数。

具体可运行代码如下:

library(Biostrings)

# 初始化并命名AAStringSet对象
df <- AAStringSet(c("MEKIVLLLA", "MEKIVLDIA"))
names(df) <- c("A1", "A2")

# 指定需要提取的位置
target_positions <- c(1,3,6,7,8,9)

# 用extractAt提取每个序列的指定位置
result <- extractAt(df, IRanges(start = target_positions, width = 1))

# 查看结果
result

运行后输出会完全符合你的期望:

width seq names

[1] 6 MKLLLA A1
[2] 6 MKLDIA A2

补充说明

  • extractAt()接受IRanges对象定义提取位置,这里IRanges(start = target_positions, width = 1)表示每个目标位置都是单个氨基酸的区间。
  • 如果需要给不同序列指定不同的提取位置,只需把位置参数改成列表格式,比如pos_list <- list(c(1,3,6,7,8,9), c(1,2,5,7,8,9)),再传入extractAt()即可。

内容的提问来源于stack exchange,提问作者star

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 07:27:14