You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何验证列表元素是否存在于Pandas DataFrame的codon列并获取位置

问题需求

验证从文件中提取的密码子列表元素是否存在于Pandas DataFrame的codon列中,并获取这些元素对应的position值。

现有实现代码

import pandas as pd
from pathlib import Path

data = pd.read_excel(r'/home/darteagam/diploma/bert/files/codon_positions.xlsx')
df = pd.DataFrame(data,columns=['position','codon','aminoacid'])
print("DataFrame Loaded!")
#print(df)

# 读取文件
with open("/home/darteagam/diploma/bert/files/bert_aa_example.txt", "r") as f1, open("/home/darteagam/diploma/bert/files/bert_nn_example.txt", "r") as f2:
    print('第31位氨基酸:')
    for aa in f1:
        prot_seq = list(aa)
        lp = len(prot_seq)
        position_aa = prot_seq[30:31]
        position_aa = list(aa[30:31]) # 验证第31位
        print(position_aa)
    for nn in f2:
        nuc_seq = nn
        x=3 
        spl=[nuc_seq[y-x:y] for y in range(x, len(nuc_seq)+x,x)]
        pos_cod = spl[30:31]
        list_codons = (list(pos_cod))
        print(list_codons)

提取得到的密码子列表输出

['ATC']
['AAC']
['ACC']
['TTT']
['GTC']
['CTC']

DataFrame输出示例

position codon aminoacid
0          1   GCT         A
1          2   GCC         A
2          3   GCA         A
3          4  GCG          A
4          5   CGT         R
..       ...   ...       ...
56        57   TAC         Y
57        58  GTT          V
58        59  GTC          V
59        60  GTA          V
60        61   GTG         V

解决方案

步骤1:统一收集提取的密码子

先修改文件读取逻辑,把所有提取到的密码子存入一个列表,避免零散打印:

import pandas as pd

# 加载DataFrame
data = pd.read_excel(r'/home/darteagam/diploma/bert/files/codon_positions.xlsx')
df = pd.DataFrame(data, columns=['position','codon','aminoacid'])
print("DataFrame加载完成!")

# 存储提取的密码子
extracted_codons = []

with open("/home/darteagam/diploma/bert/files/bert_nn_example.txt", "r") as f2:
    for nn in f2:
        nuc_seq = nn.strip()  # 移除换行符等空白字符
        if len(nuc_seq) >= 90:  # 确保序列长度足够取到第31个密码子(3*30=90位)
            x = 3 
            spl = [nuc_seq[y-x:y] for y in range(x, len(nuc_seq)+x, x)]
            pos_cod = spl[30:31]
            if pos_cod:
                extracted_codons.extend(pos_cod)

# 去重(可选,避免重复查询)
extracted_codons = list(set(extracted_codons))
print("提取的密码子列表:", extracted_codons)

步骤2:匹配密码子并获取对应position

用Pandas的isin()方法筛选匹配行,再提取对应的position值:

# 筛选匹配的行
matched_rows = df[df['codon'].isin(extracted_codons)]

# 转为字典,方便快速查询密码子对应的position
codon_pos_map = matched_rows.set_index('codon')['position'].to_dict()

# 输出验证结果
print("\n密码子匹配结果:")
for codon in extracted_codons:
    if codon in codon_pos_map:
        print(f"密码子 {codon} 对应的position:{codon_pos_map[codon]}")
    else:
        print(f"密码子 {codon}:未在DataFrame中找到")

完整整合代码

import pandas as pd

# 加载密码子位置数据
data = pd.read_excel(r'/home/darteagam/diploma/bert/files/codon_positions.xlsx')
df = pd.DataFrame(data, columns=['position','codon','aminoacid'])
print("DataFrame加载完成!")

# 提取文件中的密码子
extracted_codons = []
with open("/home/darteagam/diploma/bert/files/bert_nn_example.txt", "r") as f2:
    for nn in f2:
        nuc_seq = nn.strip()
        if len(nuc_seq) >= 90:
            spl = [nuc_seq[y-3:y] for y in range(3, len(nuc_seq)+3, 3)]
            pos_cod = spl[30:31]
            extracted_codons.extend(pos_cod)

# 去重
extracted_codons = list(set(extracted_codons))

# 匹配并获取position
matched_rows = df[df['codon'].isin(extracted_codons)]
codon_pos_map = matched_rows.set_index('codon')['position'].to_dict()

# 输出结果
print("\n验证结果:")
for codon in extracted_codons:
    pos = codon_pos_map.get(codon, "不存在")
    print(f"密码子 {codon} 对应的position:{pos}")

内容的提问来源于stack exchange,提问作者Vykov

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 08:05:22