You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于位置信息批量替换FASTA文件中指定链的字符

解决方案:用Python批量处理FASTA序列替换

步骤说明

假设你的test.fa格式为6行(每个链占2行:标题行+序列行),示例如下:

chain A
ABCDEFG
chain B
HIJKLMN
chain C
OPQRSTU

test.list每行以空格分隔,第三列对应chain A的替换规则(值为K时替换对应位置为0),第四列对应chain B(值为E时替换对应位置为1),每行对应序列的一个位置(第1行对应序列第0位,以此类推)。

完整Python脚本

def read_fasta(fasta_path):
    chains = {}
    current_chain = None
    with open(fasta_path, 'r') as f:
        for line in f:
            line = line.strip()
            if not line:
                continue
            if line.startswith('>'):
                # 提取链名(如从">chain A"中取出"A")
                current_chain = line.split()[1]
                chains[current_chain] = []
            else:
                chains[current_chain].append(line)
    # 将序列片段合并为完整字符串
    for chain in chains:
        chains[chain] = ''.join(chains[chain])
    return chains

def process_list(list_path, chains):
    # 字符串不可直接修改,转成列表操作
    chain_seqs = {k: list(v) for k, v in chains.items()}
    with open(list_path, 'r') as f:
        for idx, line in enumerate(f):
            line = line.strip()
            if not line:
                continue
            cols = line.split()
            # 处理chain A:第三列(索引2)为"K"时替换对应位置为"0"
            if len(cols) >= 3 and cols[2] == 'K':
                if idx < len(chain_seqs.get('A', [])):
                    chain_seqs['A'][idx] = '0'
            # 处理chain B:第四列(索引3)为"E"时替换对应位置为"1"
            if len(cols) >= 4 and cols[3] == 'E':
                if idx < len(chain_seqs.get('B', [])):
                    chain_seqs['B'][idx] = '1'
            # 如需添加其他链的规则,在这里扩展即可,比如:
            # if len(cols) >=5 and cols[4] == 'X':
            #     if idx < len(chain_seqs.get('C', [])):
            #         chain_seqs['C'][idx] = '2'
    # 列表转回字符串
    updated_chains = {k: ''.join(v) for k, v in chain_seqs.items()}
    return updated_chains

def write_fasta(chains, output_path):
    with open(output_path, 'w') as f:
        for chain, seq in chains.items():
            f.write(f'>chain {chain}\n')
            f.write(f'{seq}\n')

if __name__ == '__main__':
    # 根据实际文件路径修改
    fasta_files = ['test.fa', 'other.fa']  # 多个FASTA文件的列表
    list_path = 'test.list'
    
    for fasta_path in fasta_files:
        # 生成输出文件名,避免覆盖原文件
        output_path = f'modified_{fasta_path}'
        # 读取原始序列
        chains = read_fasta(fasta_path)
        # 执行替换逻辑
        updated_chains = process_list(list_path, chains)
        # 写入修改后的文件
        write_fasta(updated_chains, output_path)

使用说明

  1. 将脚本保存为modify_fasta.py
  2. 修改脚本中fasta_files和list_path为你的实际文件路径
  3. 运行脚本:python modify_fasta.py
  4. 修改后的文件会以modified_原文件名.fa的形式保存,不会覆盖原文件

关键细节

  • 脚本会自动跳过空行,避免处理无效内容
  • 加入了索引越界判断,防止test.list行数超过序列长度时出错
  • 支持批量处理多个FASTA文件,只需在fasta_files列表中添加文件名即可

内容的提问来源于stack exchange,提问作者Jalan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 09:55:19