You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Pandas开发生物信息学程序时遭遇rsID列KeyError问题

问题描述

我正在为患者报告团队开发一款生物信息学程序,用于处理患者基因检测结果:通过NCBI的rsID提取特定SNP的核苷酸数据,将其与自建参考库合并、对比,若患者核苷酸为低频率罕见位点则标记Flag。但运行脚本时,上传患者文件和群体数据后出现KeyError,提示无法找到rsID列。已知患者CSV文件存在前置表头,需跳过以获取有效数据。

代码片段
onClick('Upload Patient Files First')
patient_data = pd.read_csv(ask_path(),)

###patient_genotype = patient_data.loc[patient_data['rsID'] == rsID]['NCBI SNP Reference']
##Not using

onClick('Upload Population Frequency Data Next')
pop_ref_data = pd.read_csv(ask_path())


#Creating a dictionary of the population reference data
def pop_dict(pop_ref_data):
    pop_ref_dict = {}
    for _, row in pop_ref_data.iterrows():
        variant_data ={}
        rsID = row['rsID']
        dominant_nucleotide = row['DomNucl']
        recessive_nucleotide = row['RecNucl']
        dominant_freq = row['DomAllele']
        recessive_freq = row['RecessiveAllele']

        variant_data[dominant_nucleotide]= dominant_freq
        variant_data[recessive_nucleotide]= recessive_freq

        pop_ref_dict[rsID] = variant_data
    return pop_ref_dict
报错回溯信息

Traceback (most recent call last):
File "C:\Users\rcthu\PycharmProjects\WorkStuff\venv\lib\site-packages\pandas\core\indexes\base.py", line 3802, in get_loc
return self._engine.get_loc(casted_key)
File "pandas_libs\index.pyx", line 138, in pandas._libs.index.IndexEngine.get_loc
File "pandas_libs\index.pyx", line 165, in pandas._libs.index.IndexEngine.get_loc
File "pandas_libs\hashtable_class_helper.pxi", line 5745, in pandas._libs.hashtable.PyObjectHashTable.get_item
File "pandas_libs\hashtable_class_helper.pxi", line 5753, in pandas._libs.hashtable.PyObjectHashTable.get_item
KeyError: 'rsID'

The above exception was the direct cause of the following exception:

Traceback (most recent call last):
File "C:\Users\rcthu\AppData\Roaming\JetBrains\PyCharmCE2022.2\scratches\Flag Process 2.12.py", line 61, in
pop_ref_row = pop_dict(pop_ref_data)
File "C:\Users\rcthu\AppData\Roaming\JetBrains\PyCharmCE2022.2\scratches\Flag Process 2.12.py", line 41, in pop_dict
rsID = row['rsID']
File "C:\Users\rcthu\PycharmProjects\WorkStuff\venv\lib\site-packages\pandas\core\series.py", line 981, in getitem
return self._get_value(key)
File "C:\Users\rcthu\PycharmProjects\WorkStuff\venv\lib\site-packages\pandas\core\series.py", line 1089, in _get_value
loc = self.index.get_loc(label)
File "C:\Users\rcthu\PycharmProjects\WorkStuff\venv\lib\site-packages\pandas\core\indexes\base.py", line 3804, in get_loc
raise KeyError(key) from err
KeyError: 'rsID'

Process finished with exit code 1

解决方案

KeyError的核心原因是读取CSV时没有跳过前置表头,导致真实列名未被识别为DataFrame的列。不管是患者文件还是群体参考文件,只要存在无效前置表头,都需要在pd.read_csv中指定参数跳过对应行数,确保真实表头被正确解析。

具体修改步骤:

  1. 确认前置表头行数:打开CSV文件,统计需要跳过的无效前置行数(比如文件开头有1行说明文字,就填1)。
  2. 修改文件读取代码:
    • 患者文件读取:
      patient_data = pd.read_csv(ask_path(), skiprows=N)  # N为需要跳过的前置行数
      
    • 群体参考文件读取:
      pop_ref_data = pd.read_csv(ask_path(), skiprows=M)  # M为群体文件的前置行数
      
  3. 验证列名正确性:读取后打印列名确认,避免后续出错:
    print(patient_data.columns)
    print(pop_ref_data.columns)
    

额外建议:

  • 如果不确定前置行数,可以用header参数指定真实表头所在行号(比如真实表头在第2行,就填header=1),效果和skiprows一致。
  • 处理CSV前手动检查文件结构,确保列名与代码中引用的rsID、DomNucl等完全匹配(注意大小写、空格等细节)。

内容的提问来源于stack exchange,提问作者ClarkThark

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 02:36:31